AI models for India,up to 86% cheaperthan Sarvam.
LLMs, real-time voice AI, speech-to-text, text-to-speech and video understanding on one OpenAI-compatible API — measured against Sarvam on real 8 kHz phone calls, not a studio recording.
- OpenAI compatible
- Real 8 kHz calls
- 21 languages
- Production ready
हिन्दीमराठीதமிழ்বাংলাతెలుగుಕನ್ನಡUp to 86% lower cost
vs Sarvam
21 languages
India-first
Real-time calling
8 kHz optimized
STT • LLM • TTS
One API
Reliable & scalable
Built for production
Speech-to-text that hears 21 languages, code-mix included
See it work
The voice AI platform, built for Indian calls.
Five products, one key. Pick one, press play, and hear a call, a recording and its transcript, a voice — or watch Lark-V read a real clip.
Hear it in your language
Osprey Live
Choose a scenario
Live call · hears, thinks and speaks on one WebSocket
Illustrative call
Spoken by Pica: Tanvi is the agent, Aarav the caller. The conversation is illustrative.
नमस्ते! क्लिनिक से बोल रही हूँ। आपको किस दिन का अपॉइंटमेंट चाहिए?
कल शाम पाँच बजे के बाद मिल जाएगा?
कल साढ़े पाँच बजे का स्लॉट खाली है। बुक कर दूँ?
हाँ, कर दीजिए।
हो गया — कन्फ़र्मेशन SMS भेज दिया है।
MiniCrow vs Sarvam · every product, one table
Everything we sell, and what it actually costs next to Sarvam.
Seven capabilities, each next to Sarvam's published rate where there is one. Four cost less outright, chat depends on the tier, and two have nothing on Sarvam's side to compare.
MiniCrow
₹6.9 – 211.2 /Mtok in
3 tiers, 1M+ context, every lane
Sarvam
₹29.28 /Mtok in
Sarvam 105B · ₹73.20 out
Flash Lite and Flash cost less per token; Pro is a higher price class
MiniCrow
₹20 /hour
Lark mini, Hindi & English
Sarvam
₹30 /hour
published rate
A third less — same accuracy (82.3 vs 80.8, 102 real calls)
MiniCrow
+₹3.50 /hour
on top of the ASR tier · not billed on pre-split stereo
Sarvam
+₹15 /hour
batch with diarization ₹45 vs ₹30 without
77% less for the speaker step
MiniCrow
from ₹21 /hour
of audio · nano ₹21 · small ₹32 · large ₹60
Sarvam
≈₹151 /hour
₹30 per 10,000 chars at 14 chars a second
86% less on pica-nano
MiniCrow
₹65.74 /call-hour
all in
Sarvam
≈₹166.70 /call-hour
its own ASR + LLM + TTS, published per-unit rates
61% less — judged 81.1 vs 73.1 blind, 20.6% vs 35.4% word error
MiniCrow
Per clip
billed on what each clip used
Sarvam
No published rate
No published Sarvam equivalent
MiniCrow
₹2.1 /Mtok
dense + sparse
Sarvam
No published rate
No published Sarvam equivalent
For developers
Everything you need to add Indian-language AI to your product.
Add speech-to-text
to your app in minutes
import requests
r = requests.post(
"https://api.minicrow.com/v1/audio/transcriptions",
headers={"Authorization": "Bearer mc_YOUR_KEY"},
data={"model": "lark-mini", "language": "hi", "script": "native", "diarize": "true"},
files={"file": open("call.wav", "rb")},
)
print(r.json()["text"])OpenAI-compatible
Change the base URL and the key.
Cost per call
In paise, in every response.
The Osprey family
Four models. Pick the one that fits the job.
Three chat tiers priced per token, and Osprey Live for phone calls. Every response tells you which lane answered and what the call cost.
Osprey Flash Lite
osprey-flash-lite
The cheap one that still sees and hears.
From
₹9.5/ M tokens in
₹35.9 out
- Modes
- speed · intelligence
- Input
- text · images · audio · files
- Tools
- Not supported
- Context
- 1M+ tokens
Best for
- High-volume classification
- Tagging and routing
- Cheap extraction passes
Osprey Flash
osprey-flash
The workhorse. Nearly everything belongs here.
From
₹6.9/ M tokens in
₹19.0 out
- Modes
- speed · intelligence · max
- Input
- text · images · files
- Tools
- Function calling
- Context
- 1M+ tokens
Best for
- Chat and assistants
- Agent loops with tools
- Structured extraction
Osprey Pro
osprey-pro
A different price class, not a nicer Flash.
From
₹79.2/ M tokens in
₹396.0 out
- Modes
- speed · intelligence · max
- Input
- text · images · audio · files
- Tools
- Function calling
- Context
- 1M+ tokens
Best for
- Multi-turn agent loops
- Long-context reasoning
- Code and analysis that must be right
Osprey Live
osprey-live
The phone agent. Hears, decides and speaks on one socket.
All in
₹65.74/ call-hour
hearing, thinking and speaking
- Languages
- Hindi · English
- Input
- caller audio, voice out
- Tools
- Your own, mid-call
- Connection
- one WebSocket
Best for
- Inbound and outbound call agents
- Lead qualification
- Appointment booking
Every mode, side by side
₹ per million tokens · 1M+ context on every chat mode
| Mode | Best for | Takes | Tools | In | Out |
|---|---|---|---|---|---|
osprey-flash-liteComing soon | |||||
:speed | Thinking off. For volume. | text · images · audio · files | ₹9.5 | ₹35.9 | |
:intelligencedefault | Thinking on. The default. | text · images · audio · files | ₹9.5 | ₹35.9 | |
osprey-flash | |||||
:speed | Short exchanges, low latency. | text only | ₹6.9 | ₹19.0 | |
:intelligencedefault | The default. Tools, images, most work. | text · images | ₹7.9 | ₹26.4 | |
:max | When it has to hold up. | text · images · files | ₹21.1 | ₹126.7 | |
osprey-pro | |||||
:speed | The widest input of any lane. | text · images · audio · files | ₹79.2 | ₹396.0 | |
:intelligencedefault | The default. Deep reasoning. | text only | ₹147.8 | ₹464.6 | |
:max | The top of the catalogue. | text · images · files | ₹211.2 | ₹1,056.0 | |
osprey-liveLive | |||||
| Voice agent | Phone calls, with your own tools | caller audio · voice out | ₹65.74 / call-hour, all in | ||
Speech, video and retrieval
The rest of the catalogue, built on Indian audio.
Transcription, video summaries, speech and vector search — the same key, the same prepaid balance, the same cost in the response.
Lark
Speech to textIndian-language transcription, code-mix included. Measured on real 8 kHz call audio, not on studio recordings.
Good for
- Phone-call transcription
- Voice notes and meetings
- Hinglish and Indic speech
- Support-call QA
- Anything recorded at 8 kHz
POST /v1/audio/transcriptions
- lark-nanoComing soonShort clips, up to 5 minutes₹10per hour of audio
- lark-miniDefault lane · Hindi and English follow your sample rate; price may differ by rate and by streaming₹20per hour of audio
- lark-largeWider output budget · :max is the top mode₹30 · ₹45 · ₹60per hour — international · Indian · :max
Lark-V
VideoA summary and a seekable timeline from one upload, in three tiers by accuracy. The large tier is the most accurate — it sees the picture and hears the sound together, in the script it was spoken in.
Good for
- Video summaries
- A seekable timeline of a clip
- Screen recordings and demos
- Ad and creative review
- Spoken content inside video
POST /v1/video/summaries
- lark-v-nanoComing soonGood accuracyPer clipbilled on what it used
- lark-v-miniBetter accuracyPer clipbilled on what it used
- lark-v-largeBest accuracy — sees and hearsPer clipbilled on what it used
All three tiers are live.
Pica
Text to speechThree tiers on one endpoint, billed per second of audio. nano has eleven pinned voices and seven emotions on Hindi. small adds 25 Hindi and English voices. large is the most natural: 17 voices, each speaking Hindi and English.
Good for
- Voice notifications and IVR
- Hindi narration with emotion
- Audio versions of written content
- Product and demo voiceover
- Accessibility read-aloud
POST /v1/audio/speech
- pica-nano11 voices, 7 emotions on Hindi₹21per hour of audio
- pica-small25 Hindi and English voices₹32per hour of audio
- pica-largeThe most natural — 17 voices, Hindi and English₹60per hour of audio
Embeddings and rerank
LiveDense and sparse vectors from one call, and a cross-encoder that reranks a shortlist. Unbranded on purpose — this one has not been given a name yet.
- Search over your own documents
- RAG retrieval
- Deduplication
- Reranking a shortlist
POST /v1/embeddings · POST /v1/rerank
₹2.1
per Mtok · dense + sparse
Chat prices are per million tokens, in rupees, and move when a model is repriced. Embeddings, transcription and voice are flat list prices. Every figure on this page is read from the live price list, so a change shows here within a couple of minutes, and a change never rewrites a call already made.
The router tells you what it picked, and why.
Send mode: "auto" and rules — not a model — pick the lane in under two milliseconds. Every response names the rule, so a route you disagree with is a string you can search for.
{ "model": "osprey-flash", "x_minicrow": { "requested_mode": "auto", "served_mode": "intelligence", "route_reason": "S:tools", "cost_known": true }, "usage": { "prompt_tokens": 88, "completion_tokens": 60, "cost": 0.6458, "cost_currency": "INR_paise" }}cost is in paise, to four decimals — a short call costs a fraction of a paisa, and rounding every call up would bill a busy month wrongly.
H6:open_tool_loopNever switch mode inside an open tool loop
A tool loop is one continuous piece of reasoning: it finishes in the mode that began it, or the follow-up is refused.
S:tools← the rule behind this responseTools present, or three turns deep, takes the middle lane
A wrong cheap route breaks the agent loop. A wrong expensive one only costs more.
S:stickyA thread keeps the mode it started on
A switch throws away the prompt cache, and cache-read is a fraction of input price on every lane.
S:long_inputEight thousand tokens of input is a summarisation job
One long user turn with no tools is a different shape of work from a conversation.
H2A lane that cannot take the modality is removed
An image on a text-only lane is not a worse answer, it is an error.
EOne rung up, once, only on a verifiable failure
A tool call that cannot be executed is evidence. A truncated answer is not — that is a continuation problem in the same mode.
max is never chosen for you
The most expensive lane is reached by naming it, or by one escalation after a verifiable failure.
A mode you name is obeyed
Name a mode and it is never escalated. Billing you for a dearer lane would be an override, not a fix.
A voice agent that hears right, judged blind against Sarvam.
Osprey Live takes the caller's speech in and speaks back on one WebSocket, calling your tools in between. Measured against Sarvam's own speech-to-text, LLM and text-to-speech on the same real phone calls.
- Reply quality, blind judge81.173.1+8.0 points
- Word error on the caller's speech20.6%35.4%14.8 points lower
- Price per call-hour, all in₹65.74≈₹166.7061% less
- First audio after the caller stops2.4–3.1 s4–6 sOsprey Live faster
2 real Marathi phone calls, 58 turns — directional. The judge was never told which stack answered, and the order was shuffled twice. Prices are published list rates. Sarvam's first audio is with its LLM's thinking on (≈1.1 s with it off).
- Your own tools, mid-callDescribe the agent and its tools once; it calls them while the caller is still on the line.Live · session.configure, your own tool list
- Both languages, one socketHindi and English on the same WebSocket and the same brief.Live today
The same accuracy, for a third less.
Lark is built and measured on the audio Indian products actually have: 8 kHz phone calls, code-mixed, mostly not in English.
- Price per hour of audio₹20₹3033% less
- Speaker labels, per hour+₹3.50+₹1577% less
- Accuracy on 102 real 8 kHz calls82.380.8Level
Accuracy is level: 1.5 points on 102 calls is inside the noise. The claim is the same accuracy for less, never “more accurate”.
- Telephone audio costs nothing8 kHz calls were measured twice against wideband audio: zero accuracy lost.Measured twice · 8 kHz vs wideband
- Code-mix comes back in the script it was spoken inCode-mixed speech comes back in the script it was spoken in, and the response says which.Live · stated default
- A hint you can send, that cannot hijack the jobSend name spellings or domain terms; they are added to the instruction, never swapped for it.Live · appended, not substituted
- Indian languages, first classEvery accuracy number here was taken on Indian-language call audio.Live today
- Speaker labels with a timestamped timelinediarize=true adds who spoke when for ₹3.50 an hour, skipped and free on pre-split stereo.Live · diarize=true
ComingTranscribe and translate in one call.Not available today.
Three voice tiers, from ₹21 an hour.
Pica speaks Hindi and English. nano has seven emotions, small adds 25 voices, and large is the most natural. Every tier is billed per second of audio.
- pica-nano, per hour of audio₹21≈₹15186% less
- pica-small, per hour of audio₹32≈₹15179% less
- pica-large, most natural₹60≈₹15160% less
Sarvam publishes ₹30 per 10,000 characters; its column is that rate at 14 characters a second, about as slowly as our slower Hindi voices read. Pica is billed per second of audio received.
{
"model": "pica-nano",
"input": "…",
"voice": "ankita",
"lane": "expressive",
"emotion": "neutral"
}Seven emotions, enforced, on the expressive (Hindi) lane. The standard lane covers both languages and has none.
- Seven emotions, on the expressive laneSeven emotions on the Hindi lane; an unknown one is a 400, not a flat reading billed anyway.Live · lane=expressive
- Eleven voices, live todaySeven Hindi and four English, each pinned to a reference clip so it never drifts.Live · GET /v1/audio/voices
ComingBring your own voice: designed from a description or cloned from a clip.Not available today.
Numbers that speak for themselves.
No invented testimonials. Every card is a result we measured on real Indian call audio or a published rate, with the sample size printed under it.
Osprey Live's replies scored 81.1 against Sarvam's calling stack at 73.1 — judged blind, shuffled twice.
Osprey Live · AI voice calling2 real Marathi calls · 58 turns · directional20.6% word error on what the caller actually said, against 35.4% for Sarvam's realtime recognizer.
Osprey Live hearingSame 58 turns, same scorer82.3 against Sarvam's 80.8 on real telephone calls — level, inside the noise — for ₹20 an hour instead of ₹30.
Lark · speech to text102 real 8 kHz Marathi callsBand-limiting call audio to 8 kHz cost zero accuracy, both times it was measured. Phone audio is not the cheap end.
Lark · telephone audioMeasured twice · 8 kHz vs wideband₹21 an hour of audio on pica-nano, against about ₹151 for Sarvam's published per-character rate. ₹60 buys the most natural tier.
Pica · text to speechPublished rates · 14 characters a secondAbout ₹66 per call-hour, all in, against ≈₹167 on Sarvam's own per-unit rates.
Osprey Live · all-in pricePublished rates, same call shapeNo plans. A balance, and a rate per model.
Add rupees, call anything, and see what each call cost in its own response.
- Flash LiteComing soon₹9.5 · ₹35.9
- Flash · speed₹6.9 · ₹19
- Flash · intelligence₹7.9 · ₹26.4
- Flash · max₹21.1 · ₹126.7
- Pro · speed₹79.2 · ₹396
- Pro · intelligence₹147.8 · ₹464.6
- Pro · max₹211.2 · ₹1,056
$ at ₹96.1 to the dollar, a configured rate. Chat prices move when a model is repriced; a change never rewrites a call already made. Full price list
Build on MiniCrow and pay up to 86% less than Sarvam.
A prepaid key in rupees, one OpenAI-compatible base URL, and a response that tells you which lane answered, why, and exactly what it cost. No plans to pick and no postpaid invoice.
Or read the API docs first.
