LLM, voice & vision AIUp to 86% cheaper

AI models for India,up to 86% cheaperthan Sarvam.

LLMs, real-time voice AI, speech-to-text, text-to-speech and video understanding on one OpenAI-compatible API — measured against Sarvam on real 8 kHz phone calls, not a studio recording.

  • OpenAI compatible
  • Real 8 kHz calls
  • 21 languages
  • Production ready
The MiniCrow crow in flight over a dotted map, feathers trailing behind itहिन्दीमराठीதமிழ்বাংলাతెలుగుಕನ್ನಡ
Real-time voice AILiveLow latency · High accuracy · OpenAI compatible

Up to 86% lower cost

vs Sarvam

21 languages

India-first

Real-time calling

8 kHz optimized

STT • LLM • TTS

One API

Reliable & scalable

Built for production

Speech-to-text that hears 21 languages, code-mix included

हिन्दीमराठीবাংলাதமிழ்తెలుగుગુજરાતીಕನ್ನಡമലയാളംਪੰਜਾਬੀଓଡ଼ିଆঅসমীয়াEnglishEspañolFrançaisDeutschPortuguêsItalianoРусскийالعربية日本語한국어
21
Languages transcribed
10 Indian · Hindi · 10 international
102
Real 8 kHz calls measured
Lark vs Sarvam: 82.3 vs 80.8
20.6%
Word error on live calls
Sarvam realtime: 35.4% · 58 turns
86%
Cheaper than Sarvam
Published rates, best case

See it work

The voice AI platform, built for Indian calls.

Five products, one key. Pick one, press play, and hear a call, a recording and its transcript, a voice — or watch Lark-V read a real clip.

Hear it in your language

Osprey Live

Choose a scenario

≈₹66 / call-hour, all inYour tools, mid-callHindi · English
Call in progress

Live call · hears, thinks and speaks on one WebSocket

Illustrative call

Spoken by Pica: Tanvi is the agent, Aarav the caller. The conversation is illustrative.

नमस्ते! क्लिनिक से बोल रही हूँ। आपको किस दिन का अपॉइंटमेंट चाहिए?

कल शाम पाँच बजे के बाद मिल जाएगा?

check_appointment_slots · tomorrow, after 17:00

कल साढ़े पाँच बजे का स्लॉट खाली है। बुक कर दूँ?

हाँ, कर दीजिए।

send_sms_confirmation

हो गया — कन्फ़र्मेशन SMS भेज दिया है।

Read the voice agent API

MiniCrow vs Sarvam · every product, one table

Everything we sell, and what it actually costs next to Sarvam.

Seven capabilities, each next to Sarvam's published rate where there is one. Four cost less outright, chat depends on the tier, and two have nothing on Sarvam's side to compare.

OspreyLLM chat & tool calling/v1/chat/completions

MiniCrow

₹6.9 – 211.2 /Mtok in

3 tiers, 1M+ context, every lane

Sarvam

₹29.28 /Mtok in

Sarvam 105B · ₹73.20 out

Flash Lite and Flash cost less per token; Pro is a higher price class

LarkSpeech to text (ASR)/v1/audio/transcriptions

MiniCrow

₹20 /hour

Lark mini, Hindi & English

Sarvam

₹30 /hour

published rate

A third less — same accuracy (82.3 vs 80.8, 102 real calls)

LarkSpeaker diarizationdiarize=true

MiniCrow

+₹3.50 /hour

on top of the ASR tier · not billed on pre-split stereo

Sarvam

+₹15 /hour

batch with diarization ₹45 vs ₹30 without

77% less for the speaker step

PicaText to speech (TTS)/v1/audio/speech

MiniCrow

from ₹21 /hour

of audio · nano ₹21 · small ₹32 · large ₹60

Sarvam

≈₹151 /hour

₹30 per 10,000 chars at 14 chars a second

86% less on pica-nano

Osprey LiveReal-time AI voice agent (calling)wss://…/v1/agent/live

MiniCrow

₹65.74 /call-hour

all in

Sarvam

≈₹166.70 /call-hour

its own ASR + LLM + TTS, published per-unit rates

61% less — judged 81.1 vs 73.1 blind, 20.6% vs 35.4% word error

Lark-VVideo → summary & timeline/v1/video/summaries

MiniCrow

Per clip

billed on what each clip used

Sarvam

No published rate

No published Sarvam equivalent

Embeddings & rerank/v1/embeddings · /v1/rerank

MiniCrow

₹2.1 /Mtok

dense + sparse

Sarvam

No published rate

No published Sarvam equivalent

For developers

Everything you need to add Indian-language AI to your product.

Add speech-to-text
to your app in minutes

More
import requests

r = requests.post(
    "https://api.minicrow.com/v1/audio/transcriptions",
    headers={"Authorization": "Bearer mc_YOUR_KEY"},
    data={"model": "lark-mini", "language": "hi", "script": "native", "diarize": "true"},
    files={"file": open("call.wav", "rb")},
)
print(r.json()["text"])
Get your API key & get started

The Osprey family

Four models. Pick the one that fits the job.

Three chat tiers priced per token, and Osprey Live for phone calls. Every response tells you which lane answered and what the call cost.

Coming soon

Osprey Flash Lite

osprey-flash-lite

The cheap one that still sees and hears.

From

₹9.5/ M tokens in

₹35.9 out

Modes
speed · intelligence
Input
text · images · audio · files
Tools
Not supported
Context
1M+ tokens

Best for

  • High-volume classification
  • Tagging and routing
  • Cheap extraction passes
Live

Osprey Flash

osprey-flash

The workhorse. Nearly everything belongs here.

From

₹6.9/ M tokens in

₹19.0 out

Modes
speed · intelligence · max
Input
text · images · files
Tools
Function calling
Context
1M+ tokens

Best for

  • Chat and assistants
  • Agent loops with tools
  • Structured extraction
Live

Osprey Pro

osprey-pro

A different price class, not a nicer Flash.

From

₹79.2/ M tokens in

₹396.0 out

Modes
speed · intelligence · max
Input
text · images · audio · files
Tools
Function calling
Context
1M+ tokens

Best for

  • Multi-turn agent loops
  • Long-context reasoning
  • Code and analysis that must be right
Live

Osprey Live

osprey-live

The phone agent. Hears, decides and speaks on one socket.

All in

₹65.74/ call-hour

hearing, thinking and speaking

Languages
Hindi · English
Input
caller audio, voice out
Tools
Your own, mid-call
Connection
one WebSocket

Best for

  • Inbound and outbound call agents
  • Lead qualification
  • Appointment booking
Request access

Every mode, side by side

₹ per million tokens · 1M+ context on every chat mode

ModeBest forTakesToolsInOut
Osprey Flash Liteosprey-flash-liteComing soon
:speedThinking off. For volume.text · images · audio · files₹9.5₹35.9
:intelligencedefaultThinking on. The default.text · images · audio · files₹9.5₹35.9
Osprey Flashosprey-flash
:speedShort exchanges, low latency.text only₹6.9₹19.0
:intelligencedefaultThe default. Tools, images, most work.text · images₹7.9₹26.4
:maxWhen it has to hold up.text · images · files₹21.1₹126.7
Osprey Proosprey-pro
:speedThe widest input of any lane.text · images · audio · files₹79.2₹396.0
:intelligencedefaultThe default. Deep reasoning.text only₹147.8₹464.6
:maxThe top of the catalogue.text · images · files₹211.2₹1,056.0
Osprey Liveosprey-liveLive
Voice agentPhone calls, with your own toolscaller audio · voice out₹65.74 / call-hour, all in

Speech, video and retrieval

The rest of the catalogue, built on Indian audio.

Transcription, video summaries, speech and vector search — the same key, the same prepaid balance, the same cost in the response.

Live

Lark

Speech to text

Indian-language transcription, code-mix included. Measured on real 8 kHz call audio, not on studio recordings.

Good for

  • Phone-call transcription
  • Voice notes and meetings
  • Hinglish and Indic speech
  • Support-call QA
  • Anything recorded at 8 kHz

POST /v1/audio/transcriptions

  • lark-nanoComing soonShort clips, up to 5 minutes₹10per hour of audio
  • lark-miniDefault lane · Hindi and English follow your sample rate; price may differ by rate and by streaming₹20per hour of audio
  • lark-largeWider output budget · :max is the top mode₹30 · ₹45 · ₹60per hour — international · Indian · :max
Live

Lark-V

Video

A summary and a seekable timeline from one upload, in three tiers by accuracy. The large tier is the most accurate — it sees the picture and hears the sound together, in the script it was spoken in.

Good for

  • Video summaries
  • A seekable timeline of a clip
  • Screen recordings and demos
  • Ad and creative review
  • Spoken content inside video

POST /v1/video/summaries

  • lark-v-nanoComing soonGood accuracyPer clipbilled on what it used
  • lark-v-miniBetter accuracyPer clipbilled on what it used
  • lark-v-largeBest accuracy — sees and hearsPer clipbilled on what it used

All three tiers are live.

Live

Pica

Text to speech

Three tiers on one endpoint, billed per second of audio. nano has eleven pinned voices and seven emotions on Hindi. small adds 25 Hindi and English voices. large is the most natural: 17 voices, each speaking Hindi and English.

Good for

  • Voice notifications and IVR
  • Hindi narration with emotion
  • Audio versions of written content
  • Product and demo voiceover
  • Accessibility read-aloud

POST /v1/audio/speech

  • pica-nano11 voices, 7 emotions on Hindi₹21per hour of audio
  • pica-small25 Hindi and English voices₹32per hour of audio
  • pica-largeThe most natural — 17 voices, Hindi and English₹60per hour of audio

Embeddings and rerank

Live

Dense and sparse vectors from one call, and a cross-encoder that reranks a shortlist. Unbranded on purpose — this one has not been given a name yet.

  • Search over your own documents
  • RAG retrieval
  • Deduplication
  • Reranking a shortlist

POST /v1/embeddings · POST /v1/rerank

₹2.1

per Mtok · dense + sparse

Chat prices are per million tokens, in rupees, and move when a model is repriced. Embeddings, transcription and voice are flat list prices. Every figure on this page is read from the live price list, so a change shows here within a couple of minutes, and a change never rewrites a call already made.

The router tells you what it picked, and why.

Send mode: "auto" and rules — not a model — pick the lane in under two milliseconds. Every response names the rule, so a route you disagree with is a string you can search for.

200 · POST /v1/chat/completions
{  "model": "osprey-flash",  "x_minicrow": {    "requested_mode": "auto",    "served_mode":    "intelligence",    "route_reason":   "S:tools",    "cost_known":     true  },  "usage": {    "prompt_tokens": 88,    "completion_tokens": 60,    "cost": 0.6458,    "cost_currency": "INR_paise"  }}

cost is in paise, to four decimals — a short call costs a fraction of a paisa, and rounding every call up would bill a busy month wrongly.

  • H6:open_tool_loop

    Never switch mode inside an open tool loop

    A tool loop is one continuous piece of reasoning: it finishes in the mode that began it, or the follow-up is refused.

  • S:tools← the rule behind this response

    Tools present, or three turns deep, takes the middle lane

    A wrong cheap route breaks the agent loop. A wrong expensive one only costs more.

  • S:sticky

    A thread keeps the mode it started on

    A switch throws away the prompt cache, and cache-read is a fraction of input price on every lane.

  • S:long_input

    Eight thousand tokens of input is a summarisation job

    One long user turn with no tools is a different shape of work from a conversation.

  • H2

    A lane that cannot take the modality is removed

    An image on a text-only lane is not a worse answer, it is an error.

  • E

    One rung up, once, only on a verifiable failure

    A tool call that cannot be executed is evidence. A truncated answer is not — that is a continuation problem in the same mode.

max is never chosen for you

The most expensive lane is reached by naming it, or by one escalation after a verifiable failure.

A mode you name is obeyed

Name a mode and it is never escalated. Billing you for a dearer lane would be an override, not a fix.

A voice agent that hears right, judged blind against Sarvam.

Osprey Live takes the caller's speech in and speaks back on one WebSocket, calling your tools in between. Measured against Sarvam's own speech-to-text, LLM and text-to-speech on the same real phone calls.

Osprey LiveSarvam stack
  • Reply quality, blind judge81.1
    73.1
    +8.0 points
  • Word error on the caller's speech20.6%
    35.4%
    14.8 points lower
  • Price per call-hour, all in₹65.74
    ≈₹166.70
    61% less
  • First audio after the caller stops2.4–3.1 s4–6 sOsprey Live faster

2 real Marathi phone calls, 58 turns — directional. The judge was never told which stack answered, and the order was shuffled twice. Prices are published list rates. Sarvam's first audio is with its LLM's thinking on (≈1.1 s with it off).

  • Your own tools, mid-callDescribe the agent and its tools once; it calls them while the caller is still on the line.Live · session.configure, your own tool list
  • Both languages, one socketHindi and English on the same WebSocket and the same brief.Live today

The same accuracy, for a third less.

Lark is built and measured on the audio Indian products actually have: 8 kHz phone calls, code-mixed, mostly not in English.

Lark miniSarvam
  • Price per hour of audio₹20
    ₹30
    33% less
  • Speaker labels, per hour+₹3.50
    +₹15
    77% less
  • Accuracy on 102 real 8 kHz calls82.380.8Level

Accuracy is level: 1.5 points on 102 calls is inside the noise. The claim is the same accuracy for less, never “more accurate”.

  • Telephone audio costs nothing8 kHz calls were measured twice against wideband audio: zero accuracy lost.Measured twice · 8 kHz vs wideband
  • Code-mix comes back in the script it was spoken inCode-mixed speech comes back in the script it was spoken in, and the response says which.Live · stated default
  • A hint you can send, that cannot hijack the jobSend name spellings or domain terms; they are added to the instruction, never swapped for it.Live · appended, not substituted
  • Indian languages, first classEvery accuracy number here was taken on Indian-language call audio.Live today
  • Speaker labels with a timestamped timelinediarize=true adds who spoke when for ₹3.50 an hour, skipped and free on pre-split stereo.Live · diarize=true

ComingTranscribe and translate in one call.Not available today.

Three voice tiers, from ₹21 an hour.

Pica speaks Hindi and English. nano has seven emotions, small adds 25 voices, and large is the most natural. Every tier is billed per second of audio.

PicaSarvam
  • pica-nano, per hour of audio₹21
    ≈₹151
    86% less
  • pica-small, per hour of audio₹32
    ≈₹151
    79% less
  • pica-large, most natural₹60
    ≈₹151
    60% less

Sarvam publishes ₹30 per 10,000 characters; its column is that rate at 14 characters a second, about as slowly as our slower Hindi voices read. Pica is billed per second of audio received.

POST /v1/audio/speech
{
  "model": "pica-nano",
  "input": "…",
  "voice": "ankita",
  "lane": "expressive",
  "emotion": "neutral"
}

Seven emotions, enforced, on the expressive (Hindi) lane. The standard lane covers both languages and has none.

  • Seven emotions, on the expressive laneSeven emotions on the Hindi lane; an unknown one is a 400, not a flat reading billed anyway.Live · lane=expressive
  • Eleven voices, live todaySeven Hindi and four English, each pinned to a reference clip so it never drifts.Live · GET /v1/audio/voices

ComingBring your own voice: designed from a description or cloned from a clip.Not available today.

Measured, not quoted

Numbers that speak for themselves.

No invented testimonials. Every card is a result we measured on real Indian call audio or a published rate, with the sample size printed under it.

Measured
Osprey Live's replies scored 81.1 against Sarvam's calling stack at 73.1 — judged blind, shuffled twice.
Osprey Live · AI voice calling2 real Marathi calls · 58 turns · directional
Measured
20.6% word error on what the caller actually said, against 35.4% for Sarvam's realtime recognizer.
Osprey Live hearingSame 58 turns, same scorer
Measured
82.3 against Sarvam's 80.8 on real telephone calls — level, inside the noise — for ₹20 an hour instead of ₹30.
Lark · speech to text102 real 8 kHz Marathi calls
Measured
Band-limiting call audio to 8 kHz cost zero accuracy, both times it was measured. Phone audio is not the cheap end.
Lark · telephone audioMeasured twice · 8 kHz vs wideband
Measured
₹21 an hour of audio on pica-nano, against about ₹151 for Sarvam's published per-character rate. ₹60 buys the most natural tier.
Pica · text to speechPublished rates · 14 characters a second
Measured
About ₹66 per call-hour, all in, against ≈₹167 on Sarvam's own per-unit rates.
Osprey Live · all-in pricePublished rates, same call shape

No plans. A balance, and a rate per model.

Add rupees, call anything, and see what each call cost in its own response.

Prepaid, in rupeesNo subscription and no invoice. A key at zero gets a 402 before any model is called.
Cost in every responseusage.cost in paise, the same in a stream. If a price is unknown, nothing is charged.
Two kinds of rateChat is priced per million tokens and video per clip, on what each call used. Speech, voice and embeddings are flat rates.
Osprey — LLM chat₹ per million tokens, input · output
  • Flash LiteComing soon₹9.5 · ₹35.9
  • Flash · speed₹6.9 · ₹19
  • Flash · intelligence₹7.9 · ₹26.4
  • Flash · max₹21.1 · ₹126.7
  • Pro · speed₹79.2 · ₹396
  • Pro · intelligence₹147.8 · ₹464.6
  • Pro · max₹211.2 · ₹1,056

$ at ₹96.1 to the dollar, a configured rate. Chat prices move when a model is repriced; a change never rewrites a call already made. Full price list

Lark — speech to text₹ per hour of audio
  • lark-nano₹10
  • lark-mini₹20
  • lark-mini, live streaming₹20
  • lark-large · international, Indian, max₹30 · ₹45 · ₹60
  • Speaker labels, on top+₹3.5
Ready to cut your voice AI bill?

Build on MiniCrow and pay up to 86% less than Sarvam.

A prepaid key in rupees, one OpenAI-compatible base URL, and a response that tells you which lane answered, why, and exactly what it cost. No plans to pick and no postpaid invoice.

Start building free

Or read the API docs first.