Introducing

LarkSpeech to text for real-world India.

High-accuracy speech recognition for Indian languages — phone calls at 8 kHz, meetings, voice notes and code-mixed speech, from a file or live.

Free in the Playground today · Public launch 1 October 2026

Speech recognition is easy when the audio is clean.Real-world speech is different.

  • 8 kHz phone calls
  • Indian accents
  • Code-mixed speech
  • Several speakers
  • Background noise
  • Different scripts
  • Different languages

That is exactly where we built Lark.

You build the harness. We handle the hearing.

Building an AI caller, your own HeyPocket or Plaud, or your own NotebookLM? The part that listens is the part you shouldn't have to build. Lark hears Indian speech — phone lines, code-mix, several speakers — and hands your agent clean, structured text, so your time goes into the harness: the product, the prompts, the tools and the memory around the model.

You buildyour code

The harness

  • Your productThe app your users open — on the web, on a phone, on a device.
  • Prompts and personalityHow it thinks, what it says, how it sounds.
  • Tools and workflowsYour APIs and the actions it is allowed to take.
  • Memory and knowledgeWhat it remembers, and the sources it answers from.
  • EvalsHow you know it works, call after call.

# one turn of a voice app

audio → Lark → text

text + prompt + tools → Osprey → reply

reply → Pica → voice

Build your own HeyPocket or Plaud. Lark is the hearing inside it.

Pocket, by HeyPocket, and Plaud Note and Plaud NotePin, by Plaud, are pocket-sized AI voice recorders: they capture calls, meetings and conversations and turn them into transcripts, summaries and action items.

The recorder is the easy part. The hard part is hearing — Indian languages, code-mix, phone calls, several people talking at once. That is Lark: send each recording and get back every word and who said it; then let Osprey write the summary and the action items.

On the table
A meeting, every speaker in the transcript.
On the phone
Clipped to the back of a phone: the whole call, both sides.
On you
Conversations on the move, in the language they happen in.
  1. A recording
  2. Lark — the words and who said them
  3. Osprey — a summary and the action items
  4. Embeddings — every conversation, searchable
  5. Your app

Pocket is a product of HeyPocket; Plaud Note and Plaud NotePin are products of Plaud. They are named to describe the kind of device you can build — MiniCrow is not affiliated with them. The device in the videos is an illustration made for this page.

Build your own. An AI caller, a NotebookLM of your own.

Your own AI caller

Voice agents that answer and make phone calls — receptionists, sales and support agents, appointment and payment reminders.

Live call · 01:12

Caller: “Mera order kab tak aayega?”

track_order(8841)

Agent: “Kal shaam 6 baje tak pahunch jaayega.”

Build it on MiniCrow

Lark Live hears the caller on an 8 kHz line. Osprey decides and calls your tools. Pica speaks the reply — or run all three on one socket with Osprey Live.

  • Lark Live
  • Osprey
  • Pica
  • Osprey Live

Your own NotebookLM

Like NotebookLM, Google's AI research notebook — it answers from the sources you give it, such as documents, web pages, YouTube videos and audio, and turns them into podcast-style audio overviews.

  • Policy.pdf
  • Webinar
  • Call notes
  • Help page

Refunds are paid within 7 days of the return reaching us.2

Audio overview · 6:12

Build it on MiniCrow

Lark turns audio into text and Lark-V reads the videos. Embeddings & Reranking find the right passage, Osprey answers from it, and Pica reads it aloud.

  • Lark
  • Lark-V
  • Embeddings & Reranking
  • Osprey
  • Pica

NotebookLM is a product of Google. They are named only to describe the kind of app you can build — MiniCrow is not affiliated with them. The pictures are illustrations.

Built for how India talks. Every hard case, handled.

8 kHz

phone audio, zero accuracy lost

Telephony-grade

Measured twice against wideband: narrowing a recording to a phone line's 8 kHz cost nothing.

24

Indian languages

India's languages

Hindi, Marathi, Tamil, Telugu, Bengali, Urdu and more — with English and nine international languages.

Live

turn by turn, on one WebSocket

Real time with Lark Live

A fast draft the moment a speaker stops, and a checked final right after. In preview.

S1 · S2

who said what, and when

Speaker labels

Every turn with its speaker and timestamps — free when each speaker has their own channel.

अ · A

native script or romanised

Your script, your choice

Code-mix stays as spoken — “उद्या meeting आहे का?” — and the response says which script you got.

₹20

per hour of audio

Affordable at scale

From ₹20 an hour on lark-mini, the default — and lark-nano, the lightest, is coming soon. Billed by the second, prepaid, no subscription.

Every language. Every script.

Declare the language and choose the alphabet: each language in its own script, or romanised the way people type in chat. English words stay in English, exactly as they were said.

The Lark mascot swooping down, listening.
  • Hindiनमस्ते, मेरा ऑर्डर कब आएगा?
  • Hinglish · romanisedKal shaam 5 baje Pune me meeting hai.
  • Marathi · code-mixउद्या meeting आहे का?
  • Tamilஉங்கள் பில் தொகை எவ்வளவு?
  • Bengaliআমার অর্ডার কোথায়?

24 Indian languages also live on Lark Live

  • Hindi
  • Marathi
  • Bengali
  • Tamil
  • Telugu
  • Gujarati
  • Kannada
  • Malayalam
  • Punjabi
  • Odia
  • Assamese
  • Urdu
  • Nepali
  • Sanskrit
  • Konkani
  • Maithili
  • Sindhi
  • Kashmiri
  • Dogri
  • Manipuri
  • Santali
  • Bodo
  • Bhojpuri
  • Rajasthani

And ten more, all live

  • English
  • Spanish
  • French
  • German
  • Portuguese
  • Italian
  • Russian
  • Arabic
  • Japanese
  • Korean

Benchmarks. Measured on real Indian speech.

Lark against the three speech-to-text APIs developers compare it with, on MiniCrow's current benchmark: 8 kHz telephone calls, 24–48 kHz recordings, and how much of the meaning survives. Best result in each row in bold.

Swipe the table to compare all four

Benchmark results for Lark, Sarvam, Google Chirp 3 and ElevenLabs Scribe v2: word error rate on 8 kHz telephone audio and on 24–48 kHz recordings, semantic meaning, and list price per hour of audio.
BenchmarkLarklark-mini, by MiniCrow Most accurateSarvamSpeech to Text APIGoogleChirp 3, Cloud Speech-to-TextElevenLabsScribe v2
Accuracy
Word error rate, 8 kHz phone calls ↓120–25%18–27%22–30%22–28%
Word error rate, 24–48 kHz recordings ↓18–12%10–15%12–18%9–13%
Semantic meaning ↑295%91%90%90%
Price
List price, per hour of audio3₹20$0.208batch and real time (Lark Live)₹30$0.312≈₹92$0.96≈₹37$0.39 · real timebatch ≈₹21 · $0.22
  1. Word error rate: the share of words transcribed wrongly — lower is better. Ranges across languages and test sets; bold marks the lowest upper bound.
  2. Semantic meaning: how much of what the speaker meant survives in the transcript, scored against the script — higher is better.
  3. Published list prices on 24 September 2026, rupees at ₹96.1 to the dollar: Lark is lark-mini, live from our price book; Sarvam's Speech to Text API ₹30 an hour; Google Cloud Speech-to-Text with Chirp 3, $0.016 a minute; ElevenLabs Scribe v2 on its developer API, $0.22 an hour and $0.39 in real time (lowered in May 2026 from $0.40; its creator app bills in credits instead). Check each provider for current rates.

MiniCrow's current benchmark measurements; results vary by language, accent, recording quality, dataset and evaluation methodology. Sarvam, Google and ElevenLabs are third-party providers named for comparison only; MiniCrow is not affiliated with them.

Word error rate, 8 kHz phone callsLower is better — the range across languages and test sets.
Lark
20–25
Sarvam
18–27
Google
22–30
ElevenLabs
22–28
Word error rate, 24–48 kHz recordingsLower is better.
Lark
8–12
Sarvam
10–15
Google
12–18
ElevenLabs
9–13
95%

95% of the meaning, kept

More than just transcription: Lark understands Indian voices, accents and context.

Measured · 102 real calls

A third less, at the same accuracy

Lark mini sells at ₹20 per hour of audio. Sarvam's published rate is ₹30. On 102 real 8 kHz Marathi calls the two scored 82.3 and 80.8 — level, inside the noise. You are not trading accuracy for the price.

Hear it. Read it. Real recordings, real Lark transcripts.

Press play: a support call, a voice note and a meeting, voiced by Pica and transcribed by lark-mini — word for word, as the audio reaches them.

Open the Playground

Hear it in your language

Lark

Choose a scenario

₹20 / hour+₹3.50 / hour for speakersBatch or live

8 kHz telephone audio · 21 languages

Lark's transcript

Spoken in Hindi by Pica's Aarav and Trisha; the words are lark-mini's own transcript, timed to the recording.

Speaker 100:00.3

Hello, main aapke order ke baare mein call kar raha hoon.

Speaker 200:03.4

Haan ji, order abhi tak deliver nahi hua, tracking bhi update nahi ho rahi.

Speaker 100:08.1

Main check karta hoon, ek minute hold kariye.

lark-mini · language=hi · diarize=true · script=latin
Read the transcription API

Wherever people talk. Lark is listening.

Contact centres
Every customer call transcribed as it happens — for summaries, quality checks and agent assist.
Meetings
A recorder on the table, a transcript with its speakers, and the action items after.
  • AI calling & contact centresSupport and sales calls, summaries, quality monitoring and agent assist.
  • Voice AI agentsThe hearing for AI receptionists, sales, support and appointment agents.
  • Meetings & note takingTranscripts with speakers, summaries, action items, searchable talks.
  • EducationLectures and classes, across Indian languages.
  • Content creationPodcasts, interviews and videos, turned into searchable text.
  • Multilingual appsApps that understand speech in the language people actually speak.
  • Enterprise AISearchable voice archives and conversation intelligence.
  • AccessibilityCaptions and transcripts for anyone who can't hear the audio.
  • AI companionsCompanions that listen in your own language.
  • Developer toolsOpenAI's request shape, with the cost in every response.

Illustrative scenes.

More than speech to text. Audio, then meaning, then action.

  1. Human speech
  2. Lark
  3. Osprey
  4. Pica
  5. Human

Lark handles the hearing. Osprey handles the reasoning. Pica handles the voice. And Osprey Live brings the real-time experience together — one platform instead of a dozen services stitched together.

A recording

One request: the language, the script, and who spoke when.

Python

import requests

r = requests.post(
    "https://api.minicrow.com/v1/audio/transcriptions",
    headers={"Authorization": "Bearer mc_YOUR_KEY"},
    data={
        "model": "lark-mini",
        "language": "hi",    # declared, never guessed
        "script": "latin",   # or "native"
        "diarize": "true",   # who spoke when
        "speakers": "2",
    },
    files={"file": open("call.wav", "rb")},
)
print(r.json()["text"])

A live call

Stream the line in; read every turn back as the caller speaks.

Python

# pip install "websockets>=14"
import asyncio, json, websockets

URL = ("wss://api.minicrow.com/v1/audio/transcriptions/live"
       "?model=lark-mini&language=mr&script=latin"
       "&encoding=mulaw&sample_rate=8000")
KEY = {"Authorization": "Bearer mc_YOUR_KEY"}

async def main():
    async with websockets.connect(URL, additional_headers=KEY) as ws:
        # stream the call's audio in as binary frames
        async for message in ws:
            event = json.loads(message)
            if event["type"] == "transcript.final":
                print(event["text"])

asyncio.run(main())
  • Declare the language

    Lark listens for the language you name — it never guesses.

  • Choose the alphabet

    script=native or script=latin — English words stay in English.

  • Who spoke when

    diarize=true adds speaker turns with start and end times, beside an unchanged transcript.

  • Two channels, two speakers

    A stereo call with one party per channel is transcribed channel by channel — no labels to buy.

  • Names and terms

    domain, vocabulary and abbreviations help Lark spell the names in your recordings.

  • The formats you have

    WAV, OGG/Opus, MP3, M4A or WebM, up to 25 MB — billed by the second, read from the file's own header.

  • Cost in every response

    In paise, on every call — and nothing charged when a transcription fails.

  • Live, over one WebSocket

    μ-law, A-law or PCM at 8 or 16 kHz; a draft and a checked final for every turn.

Every field, limit and error is in the API reference.

Which Lark is right for you?

Three tiers, one endpoint. Live prices per hour of audio, billed by the second, prepaid in rupees.

The Lark mascot, a glowing blue and magenta lark, swooping down to listen.

Coming soon

Lark nano

The lightest Lark.

The lightest Lark — for short clips and voice notes at volume.

₹10

per hour of audio · $0.104 an hour

Coming soonDocs
Model id
lark-nano
Limit
5 minutes of audio a request
Real time
Upload only
Best for
Voice notes · short clips · volume
The Lark mascot, a glowing blue and magenta lark, face on with its wings raised.

The default

Lark mini

The one to build on.

The default — for calls, meetings and voice notes, in every language.

₹20

per hour of audio · $0.208 an hour

Model id
lark-mini
Limit
25 MB a request
Real time
Yes — Lark Live
Best for
Calls · meetings · voice agents
The Lark mascot seen from above, wings spread.

 

Lark large

For audio that must be right.

For the hardest audio — and lark-large:max, the most accurate Lark.

₹30–₹60

₹30 · ₹45 Indian · ₹60 :max

Model id
lark-large
Limit
25 MB a request
Real time
Upload only
Best for
Hard audio · long meetings · archives

Lark Live Preview

Real-time speech to text over a WebSocket — ₹20 an hour ($0.208), billed per second of audio you send. Enabled per account.

Speaker labels

+₹3.50 an hour on any tier ($0.036) — not charged if the speaker pass fails, and free on a two-channel call.

Every price
ModelForPer hour
lark-nanocoming soonShort clips, up to 5 minutes a request₹10$0.104
lark-miniThe default — uploads and Lark Live₹20$0.208
lark-largeNo Indian language declared₹30$0.312
lark-largeAn Indian language declared₹45$0.468
lark-large:maxThe most accurate Lark₹60$0.624
Speaker labelsdiarize=true, on any tier; free on a two-channel call+₹3.50$0.036

Live from our price book; dollars at ₹96.1. Every MiniCrow price

Production-grade speech recognitionshouldn't require production-grade pricing.

  • Affordable.
  • Multilingual.
  • Real-time.
  • Built for real conversations.

Public launch

1 October 2026.

Build › Deploy › Scale

High-quality speech intelligence, for every developer. Try Lark free in the Playground today — no account needed.

Two Lark mascots in flight, one singing to the other.

Questions.

What is Lark?

Lark is MiniCrow's speech-to-text (STT, or ASR) API for Indian languages and real-world audio — 8 kHz phone calls, Indian accents, code-mixed speech, several speakers and background noise. It turns speech into clean, structured text for AI voice agents, contact centres, meetings, transcription products and multilingual apps.

Which languages does Lark support?

24 Indian languages on upload — Hindi, Marathi, Bengali, Tamil, Telugu, Gujarati, Kannada, Malayalam, Punjabi, Odia, Assamese, Urdu, Nepali, Sanskrit, Konkani, Maithili, Sindhi, Kashmiri, Dogri, Manipuri, Santali, Bodo, Bhojpuri, Rajasthani — and English, Spanish, French, German, Portuguese, Italian, Russian, Arabic, Japanese, Korean. Lark Live, the real-time stream, covers 11 of the Indian languages and all ten international ones. Declare the language with each request: Lark never guesses. Accuracy is measured most on Hindi and Marathi phone calls, and on read speech in 14 languages.

Does Lark work on 8 kHz phone calls?

Yes — it is built for them. Telephone audio at 8 kHz was measured against wideband recordings twice, and narrowing the band cost zero accuracy both times. Send μ-law, A-law or PCM straight from the line; there is no need to upsample.

Can Lark tell who is speaking?

Yes. Send diarize=true — and speakers=2 when you know it — and every turn comes back with its speaker and its start and end times, beside an unchanged transcript. It adds ₹3.50 an hour to the tier's price, and nothing when the speaker pass fails or when a stereo call already carries one speaker per channel.

Can I get romanised text instead of the native script?

Yes, it is your choice: script=native writes each language in its own script, script=latin romanises it the way people type in chat. Code-mix stays as spoken — English words stay in English — and the response says which script you got.

Does Lark work in real time?

Yes, with Lark Live: a WebSocket that answers every turn twice — a fast draft the moment the speaker stops, and a checked final right after. It costs ₹20 an hour, billed per second. Lark Live is in preview and enabled per account.

Can Lark translate speech?

Speech translation is coming to Lark; it is not part of the API yet. Lark-V can already write a video's summary in another language — a Tamil video summarised in English, for example.

How much does Lark cost?

lark-mini ₹20 an hour, lark-large ₹30 (₹45 for an Indian language, ₹60 on lark-large:max) — $0.208 and $0.312–$0.624. lark-nano, the lightest tier, is coming soon. Billed per second of audio, prepaid in rupees, no subscription. Speaker labels add ₹3.50 an hour.

How do I build my own HeyPocket, Plaud or AI caller with Lark?

Lark is the hearing; you build the harness around it. For an AI recorder like HeyPocket's Pocket or Plaud Note, send each recording to Lark with speaker labels, then ask Osprey for the summary and the action items. For an AI caller, stream the call to Lark Live, let Osprey decide and call your tools, and answer with Pica — or use Osprey Live, which runs all three on one socket.

Is Lark OpenAI-compatible?

The request is OpenAI's shape — POST /v1/audio/transcriptions, multipart, with a file and a model — plus MiniCrow's own fields for language, script and speakers. Every response carries its cost.

When does Lark launch?

MiniCrow launches Lark publicly on 1 October 2026. You can try it free in the Playground today.

Meet the MiniCrow family. One API that talks, hears, speaks and watches.

Lark hears; the rest of MiniCrow thinks, speaks and watches. Press play: hear a real call, a voice — or watch Lark-V read a real clip.

Osprey

Thinks

Human-like conversational AI — fast, affordable LLMs for assistants, companions and agents, with tool calling.

  • LLM for real-world applications
  • 3 models · 8 modes

Plan my Sunday — something relaxed.

Brunch at 11, a walk by the lake at 4. Shall I book the table?

₹6.9

per million input tokens · $0.072

Pica

Speaks

Natural text to speech in Hindi and English, with expressive voices for assistants, IVR and narration.

  • text-to-speech API for Hindi and English
  • 3 tiers · nano · small · large

“Welcome back, Riya! Your order is on its way.”

₹21

per hour of audio · $0.219

Lark-V

Watches

Video understanding — a summary, a seekable timeline and the spoken words, from a single clip.

  • video summary and timeline API
  • 3 tiers · nano · mini · large
  • 00:04A car pulls into the driveway
  • 00:12Two people unload boxes
  • 00:31Someone waves at the camera

Per clip

billed on what each clip used

Osprey Live

Talks

Real-time multimodal AI that sees, hears, thinks and talks — with tool calling, for AI calling, companions, robots, glasses and cars.

  • real-time multimodal AI model
Live call · 00:42

Caller: “Do you deliver on Sundays?”

Osprey Live: “Yes — until 9 pm. Shall I book you a slot?”

₹65.74

per call-hour, all in · $0.684

Embeddings & Reranking Finds

Semantic search and RAG — embed your documents, rerank a shortlist, and give Osprey your own knowledge to answer from.

Docs