Free in the Playground today · Public launch 1 October 2026
Speech recognition is easy when the audio is clean.Real-world speech is different.
8 kHz phone calls
Indian accents
Code-mixed speech
Several speakers
Background noise
Different scripts
Different languages
That is exactly where we built Lark.
You build the harness. We handle the hearing.
Building an AI caller, your own HeyPocket or Plaud, or your own NotebookLM? The part that listens is the part you shouldn't have to build. Lark hears Indian speech — phone lines, code-mix, several speakers — and hands your agent clean, structured text, so your time goes into the harness: the product, the prompts, the tools and the memory around the model.
You buildyour code
The harness
Your productThe app your users open — on the web, on a phone, on a device.
Prompts and personalityHow it thinks, what it says, how it sounds.
Tools and workflowsYour APIs and the actions it is allowed to take.
Memory and knowledgeWhat it remembers, and the sources it answers from.
One keyOne billPrepaid in rupeesCost in every response
Build your own HeyPocket or Plaud. Lark is the hearing inside it.
Pocket, by HeyPocket, and Plaud Note and Plaud NotePin, by Plaud, are pocket-sized AI voice recorders: they capture calls, meetings and conversations and turn them into transcripts, summaries and action items.
The recorder is the easy part. The hard part is hearing — Indian languages, code-mix, phone calls, several people talking at once. That is Lark: send each recording and get back every word and who said it; then let Osprey write the summary and the action items.
On the table
A meeting, every speaker in the transcript.
On the phone
Clipped to the back of a phone: the whole call, both sides.
On you
Conversations on the move, in the language they happen in.
A recording
Lark — the words and who said them
Osprey — a summary and the action items
Embeddings — every conversation, searchable
Your app
Pocket is a product of HeyPocket; Plaud Note and Plaud NotePin are products of Plaud. They are named to describe the kind of device you can build — MiniCrow is not affiliated with them. The device in the videos is an illustration made for this page.
Build your own. An AI caller, a NotebookLM of your own.
Your own AI caller
Voice agents that answer and make phone calls — receptionists, sales and support agents, appointment and payment reminders.
Live call · 01:12
Caller: “Mera order kab tak aayega?”
track_order(8841)
Agent: “Kal shaam 6 baje tak pahunch jaayega.”
Build it on MiniCrow
Lark Live hears the caller on an 8 kHz line. Osprey decides and calls your tools. Pica speaks the reply — or run all three on one socket with Osprey Live.
Lark Live
Osprey
Pica
Osprey Live
Your own NotebookLM
Like NotebookLM, Google's AI research notebook — it answers from the sources you give it, such as documents, web pages, YouTube videos and audio, and turns them into podcast-style audio overviews.
Policy.pdf
Webinar
Call notes
Help page
Refunds are paid within 7 days of the return reaching us.2
Audio overview · 6:12
Build it on MiniCrow
Lark turns audio into text and Lark-V reads the videos. Embeddings & Reranking find the right passage, Osprey answers from it, and Pica reads it aloud.
Lark
Lark-V
Embeddings & Reranking
Osprey
Pica
NotebookLM is a product of Google. They are named only to describe the kind of app you can build — MiniCrow is not affiliated with them. The pictures are illustrations.
Built for how India talks. Every hard case, handled.
8 kHz
phone audio, zero accuracy lost
Telephony-grade
Measured twice against wideband: narrowing a recording to a phone line's 8 kHz cost nothing.
24
Indian languages
India's languages
Hindi, Marathi, Tamil, Telugu, Bengali, Urdu and more — with English and nine international languages.
Live
turn by turn, on one WebSocket
Real time with Lark Live
A fast draft the moment a speaker stops, and a checked final right after. In preview.
S1 · S2
who said what, and when
Speaker labels
Every turn with its speaker and timestamps — free when each speaker has their own channel.
अ · A
native script or romanised
Your script, your choice
Code-mix stays as spoken — “उद्या meeting आहे का?” — and the response says which script you got.
₹20
per hour of audio
Affordable at scale
From ₹20 an hour on lark-mini, the default — and lark-nano, the lightest, is coming soon. Billed by the second, prepaid, no subscription.
Every language. Every script.
Declare the language and choose the alphabet: each language in its own script, or romanised the way people type in chat. English words stay in English, exactly as they were said.
Hindiनमस्ते, मेरा ऑर्डर कब आएगा?
Hinglish · romanisedKal shaam 5 baje Pune me meeting hai.
Marathi · code-mixउद्या meeting आहे का?
Tamilஉங்கள் பில் தொகை எவ்வளவு?
Bengaliআমার অর্ডার কোথায়?
“Every word, as spoken.”
24 Indian languages also live on Lark Live
Hindi
Marathi
Bengali
Tamil
Telugu
Gujarati
Kannada
Malayalam
Punjabi
Odia
Assamese
Urdu
Nepali
Sanskrit
Konkani
Maithili
Sindhi
Kashmiri
Dogri
Manipuri
Santali
Bodo
Bhojpuri
Rajasthani
And ten more, all live
English
Spanish
French
German
Portuguese
Italian
Russian
Arabic
Japanese
Korean
Benchmarks. Measured on real Indian speech.
Lark against the three speech-to-text APIs developers compare it with, on MiniCrow's current benchmark: 8 kHz telephone calls, 24–48 kHz recordings, and how much of the meaning survives. Best result in each row in bold.
Swipe the table to compare all four
Benchmark results for Lark, Sarvam, Google Chirp 3 and ElevenLabs Scribe v2: word error rate on 8 kHz telephone audio and on 24–48 kHz recordings, semantic meaning, and list price per hour of audio.
Benchmark
Larklark-mini, by MiniCrow Most accurate
SarvamSpeech to Text API
GoogleChirp 3, Cloud Speech-to-Text
ElevenLabsScribe v2
Accuracy
Word error rate, 8 kHz phone calls ↓1
20–25%
18–27%
22–30%
22–28%
Word error rate, 24–48 kHz recordings ↓1
8–12%
10–15%
12–18%
9–13%
Semantic meaning ↑2
95%
91%
90%
90%
Price
List price, per hour of audio3
₹20$0.208batch and real time (Lark Live)
₹30$0.312
≈₹92$0.96
≈₹37$0.39 · real timebatch ≈₹21 · $0.22
Word error rate: the share of words transcribed wrongly — lower is better. Ranges across languages and test sets; bold marks the lowest upper bound.
Semantic meaning: how much of what the speaker meant survives in the transcript, scored against the script — higher is better.
Published list prices on 24 September 2026, rupees at ₹96.1 to the dollar: Lark is lark-mini, live from our price book; Sarvam's Speech to Text API ₹30 an hour; Google Cloud Speech-to-Text with Chirp 3, $0.016 a minute; ElevenLabs Scribe v2 on its developer API, $0.22 an hour and $0.39 in real time (lowered in May 2026 from $0.40; its creator app bills in credits instead). Check each provider for current rates.
MiniCrow's current benchmark measurements; results vary by language, accent, recording quality, dataset and evaluation methodology. Sarvam, Google and ElevenLabs are third-party providers named for comparison only; MiniCrow is not affiliated with them.
Word error rate, 8 kHz phone callsLower is better — the range across languages and test sets.
Lark
20–25
Sarvam
18–27
Google
22–30
ElevenLabs
22–28
0102030
Word error rate, 24–48 kHz recordingsLower is better.
Lark
8–12
Sarvam
10–15
Google
12–18
ElevenLabs
9–13
0102030
95% of the meaning, kept
More than just transcription: Lark understands Indian voices, accents and context.
Measured · 102 real calls
A third less, at the same accuracy
Lark mini sells at ₹20 per hour of audio. Sarvam's published rate is ₹30. On 102 real 8 kHz Marathi calls the two scored 82.3 and 80.8 — level, inside the noise. You are not trading accuracy for the price.
Hear it. Read it. Real recordings, real Lark transcripts.
Press play: a support call, a voice note and a meeting, voiced by Pica and transcribed by lark-mini — word for word, as the audio reaches them.
EducationLectures and classes, across Indian languages.
Content creationPodcasts, interviews and videos, turned into searchable text.
Multilingual appsApps that understand speech in the language people actually speak.
Enterprise AISearchable voice archives and conversation intelligence.
AccessibilityCaptions and transcripts for anyone who can't hear the audio.
AI companionsCompanions that listen in your own language.
Developer toolsOpenAI's request shape, with the cost in every response.
Illustrative scenes.
More than speech to text. Audio, then meaning, then action.
Human speech
Lark
Osprey
Pica
Human
Lark handles the hearing. Osprey handles the reasoning. Pica handles the voice. And Osprey Live brings the real-time experience together — one platform instead of a dozen services stitched together.
A recording
One request: the language, the script, and who spoke when.
Python
import requests
r = requests.post(
"https://api.minicrow.com/v1/audio/transcriptions",
headers={"Authorization": "Bearer mc_YOUR_KEY"},
data={
"model": "lark-mini",
"language": "hi", # declared, never guessed"script": "latin", # or "native""diarize": "true", # who spoke when"speakers": "2",
},
files={"file": open("call.wav", "rb")},
)
print(r.json()["text"])
A live call
Stream the line in; read every turn back as the caller speaks.
Python
# pip install "websockets>=14"
import asyncio, json, websockets
URL = ("wss://api.minicrow.com/v1/audio/transcriptions/live""?model=lark-mini&language=mr&script=latin""&encoding=mulaw&sample_rate=8000")
KEY = {"Authorization": "Bearer mc_YOUR_KEY"}
async def main():
async with websockets.connect(URL, additional_headers=KEY) as ws:
# stream the call's audio in as binary frames
async for message in ws:
event = json.loads(message)
if event["type"] == "transcript.final":
print(event["text"])
asyncio.run(main())
Declare the language
Lark listens for the language you name — it never guesses.
Choose the alphabet
script=native or script=latin — English words stay in English.
Who spoke when
diarize=true adds speaker turns with start and end times, beside an unchanged transcript.
Two channels, two speakers
A stereo call with one party per channel is transcribed channel by channel — no labels to buy.
Names and terms
domain, vocabulary and abbreviations help Lark spell the names in your recordings.
The formats you have
WAV, OGG/Opus, MP3, M4A or WebM, up to 25 MB — billed by the second, read from the file's own header.
Cost in every response
In paise, on every call — and nothing charged when a transcription fails.
Live, over one WebSocket
μ-law, A-law or PCM at 8 or 16 kHz; a draft and a checked final for every turn.
Every field, limit and error is in the API reference.
Which Lark is right for you?
Three tiers, one endpoint. Live prices per hour of audio, billed by the second, prepaid in rupees.
Coming soon
Lark nano
The lightest Lark.
The lightest Lark — for short clips and voice notes at volume.
Lark is MiniCrow's speech-to-text (STT, or ASR) API for Indian languages and real-world audio — 8 kHz phone calls, Indian accents, code-mixed speech, several speakers and background noise. It turns speech into clean, structured text for AI voice agents, contact centres, meetings, transcription products and multilingual apps.
Which languages does Lark support?
24 Indian languages on upload — Hindi, Marathi, Bengali, Tamil, Telugu, Gujarati, Kannada, Malayalam, Punjabi, Odia, Assamese, Urdu, Nepali, Sanskrit, Konkani, Maithili, Sindhi, Kashmiri, Dogri, Manipuri, Santali, Bodo, Bhojpuri, Rajasthani — and English, Spanish, French, German, Portuguese, Italian, Russian, Arabic, Japanese, Korean. Lark Live, the real-time stream, covers 11 of the Indian languages and all ten international ones. Declare the language with each request: Lark never guesses. Accuracy is measured most on Hindi and Marathi phone calls, and on read speech in 14 languages.
Does Lark work on 8 kHz phone calls?
Yes — it is built for them. Telephone audio at 8 kHz was measured against wideband recordings twice, and narrowing the band cost zero accuracy both times. Send μ-law, A-law or PCM straight from the line; there is no need to upsample.
Can Lark tell who is speaking?
Yes. Send diarize=true — and speakers=2 when you know it — and every turn comes back with its speaker and its start and end times, beside an unchanged transcript. It adds ₹3.50 an hour to the tier's price, and nothing when the speaker pass fails or when a stereo call already carries one speaker per channel.
Can I get romanised text instead of the native script?
Yes, it is your choice: script=native writes each language in its own script, script=latin romanises it the way people type in chat. Code-mix stays as spoken — English words stay in English — and the response says which script you got.
Does Lark work in real time?
Yes, with Lark Live: a WebSocket that answers every turn twice — a fast draft the moment the speaker stops, and a checked final right after. It costs ₹20 an hour, billed per second. Lark Live is in preview and enabled per account.
Can Lark translate speech?
Speech translation is coming to Lark; it is not part of the API yet. Lark-V can already write a video's summary in another language — a Tamil video summarised in English, for example.
How much does Lark cost?
lark-mini ₹20 an hour, lark-large ₹30 (₹45 for an Indian language, ₹60 on lark-large:max) — $0.208 and $0.312–$0.624. lark-nano, the lightest tier, is coming soon. Billed per second of audio, prepaid in rupees, no subscription. Speaker labels add ₹3.50 an hour.
How do I build my own HeyPocket, Plaud or AI caller with Lark?
Lark is the hearing; you build the harness around it. For an AI recorder like HeyPocket's Pocket or Plaud Note, send each recording to Lark with speaker labels, then ask Osprey for the summary and the action items. For an AI caller, stream the call to Lark Live, let Osprey decide and call your tools, and answer with Pica — or use Osprey Live, which runs all three on one socket.
Is Lark OpenAI-compatible?
The request is OpenAI's shape — POST /v1/audio/transcriptions, multipart, with a file and a model — plus MiniCrow's own fields for language, script and speakers. Every response carries its cost.
When does Lark launch?
MiniCrow launches Lark publicly on 1 October 2026. You can try it free in the Playground today.
Meet the MiniCrow family. One API that talks, hears, speaks and watches.
Lark hears; the rest of MiniCrow thinks, speaks and watches. Press play: hear a real call, a voice — or watch Lark-V read a real clip.
Osprey
Thinks
Human-like conversational AI — fast, affordable LLMs for assistants, companions and agents, with tool calling.
LLM for real-world applications
3 models · 8 modes
Plan my Sunday — something relaxed.
Brunch at 11, a walk by the lake at 4. Shall I book the table?