MiniCrow API
Base URL https://api.minicrow.com
Compatibility OpenAI Chat Completions. An existing OpenAI client works by changing two things: the base URL
and the key.
Everything on this page is live and was verified against the deployed service, not against a mock.
For AI agents and quick integration: /agent.md is the short, prescriptive version of this reference —
every model, when to call it, the request shape, and the rules that decide accuracy (docs/AGENT.md in the repo).
Authentication
Authorization: Bearer mc_<prefix>_<secret>
A key looks like mc_96bf0550e045_kQ3v…. The middle segment is the prefix — it identifies the key in a
listing or a log line and is not secret. The whole string is the secret and is shown once, at creation, and
never again: only an argon2id hash of it is stored, so a lost key is replaced, not recovered.
Every failure to authenticate answers the same 401 invalid_api_key with the same message, whether the key does
not exist or the secret is wrong. A different answer for each would tell an attacker which half to keep guessing.
Credits and keys
Credits belong to the account. Keys draw on them. Make as many keys as you like — one per service, one per
environment — and they all spend the same balance. A new key works immediately; there is nothing to top up on it.
A key may carry its own spend limit. That is what a key for a script, a contractor or a staging environment
is for: it stops at its own cap and the rest of your balance stays usable.
|
|
402 insufficient_credit |
the account is out of credit |
402 key_limit_reached |
this key has hit its own limit; the account has not |
403 account_suspended |
the account is suspended. Not a credential problem — rotating the key will not help |
Both are checked before the model is called, so neither costs you anything.
⚠️ A key with no limit set has no limit — it is not a limit of zero. A key created without one works.
⚠️ A key's limit is a LIFETIME limit. Nothing resets it: spent counts everything that key has ever spent,
so a cap of ₹500 is ₹500 for the life of the key, not ₹500 a month. There is no billing period in the service
yet. Raise the cap, or mint a new key, when it is reached.
POST /v1/chat/completions
Standard OpenAI request body. Two things differ, and both are additions rather than changes.
Every MiniCrow model presents itself as MiniCrow's
0.42.0. A generative model here answers as the product it was called as — osprey-flash, osprey-pro, or the
persona your Osprey Live spec gives it — and as MiniCrow's. Transcription and summarising models (Lark, Lark Live,
Lark-V) produce only a transcript or a timeline and are not asked such questions.
The model id carries the mode
| You send |
You get |
osprey-flash |
the tier's default lane. Reported as default:<mode>, not as a request you made. |
osprey-flash:speed |
that lane, always. An explicit mode is a hard choice, not a hint. |
osprey-flash:intelligence |
" |
osprey-flash:max |
" |
osprey-flash:auto |
MiniCrow picks the mode for each request, from the request itself (tools, length, script). |
GET /v1/models lists every tier and the modes it offers: what you buy is the lane.
Every response says which lane answered and why
"x_minicrow": {
"requested_mode": "auto",
"served_mode": "intelligence",
"route_reason": "S:tools",
"cost_known": true
}
This is deliberate and is not going away: a choice that is hidden cannot be debugged, so the lane MiniCrow
picked and the rule that picked it are in every response.
⚠️ The lane is the product. You buy a lane at a price; the lane is what you can act on.
route_reason distinguishes explicit:<mode> (you named it), default:<mode> (you sent a bare model id) and
MiniCrow's own reasons for the mode it picked, such as S:tools.
Cost is in the response, in paise
"usage": {
"prompt_tokens": 88, "completion_tokens": 60,
"cost": 0.6458,
"cost_currency": "INR_paise"
}
⚠️ usage.cost is always what MiniCrow charges you, and it is the same number in a
streamed response as in a non-streamed one. A typical short call costs a fraction of a paisa, which is why the
field is decimal and why balances are held to six decimal places — rounding each call to a whole paisa would
report a busy month as free.
cost_known: false means the call's cost could not be measured and the lane has no flat rate, so nothing was
charged. That is a gap you should see, not a discount.
Thinking, and the room it needs
Some modes think before they answer. Thinking tokens count toward max_tokens and are billed as output; the
number is in usage.completion_tokens_details.reasoning_tokens. If the budget runs out while the model is still
thinking, the answer comes back empty with finish_reason: "length" — and is billed. Give a thinking mode room
(2,000 tokens or more for a long input), or switch thinking off where the task does not need it.
| Mode |
Thinks |
reasoning_effort |
speed |
yes, by default |
low · high · max set how much; none (or minimal, or "reasoning": {"enabled": false}) switches it off for this request |
intelligence |
always |
low · high · max set how much; it cannot be switched off |
max |
yes |
minimal · low · medium · high · max |
Any other value is ignored. Measured on osprey-flash:speed over 60 requests: with thinking off it answered in
3.0 s against 7.9 s (median), never came back empty (5 of 60 did with thinking at max_tokens 2,000, and 23 of 60
would have at 400), returned valid JSON 30 of 30 times against 24, and was judged better on call summaries — but
worse on classifications, chat replies and rewrites, and got 5 of 10 calculations wrong. Switch it off for
summaries and extraction; keep it for anything that has to be worked out.
Streaming
"stream": true returns text/event-stream, terminated by data: [DONE].
Every chunk carries the model id you called, and the final usage chunk carries usage.cost in paise plus
x_minicrow. Nothing else about how the lane was chosen is included. Cloudflare
does not buffer these streams; chunks were measured arriving 20–30 ms apart through the public host.
Each event is written and flushed on its own, the moment it arrives: one data: {…}\n\n per event.
A failure that happens before the first event — a lane that is down, a rate limit, a model that accepted the
call and then said nothing — is not a stream at all: it is the ordinary JSON error envelope with its own status,
exactly as without stream, and it is not charged.
A stream that breaks after it started ends on an error event, in the same envelope as every other error:
data: {"error":{"message":"The answer was cut off before it was complete. Try again.","type":"api_error","code":"model_unavailable"}}
It is the last event, and there is no data: [DONE] after it — so an OpenAI client raises instead of reading
the partial answer as a whole one. When the model itself reports a failure mid-answer, that chunk keeps its
place with an error object in the same shape (model_unavailable, rate_limited, model_timeout, …).
A stream that breaks part-way is still charged for what was produced: a partial answer you received is not free
to make. Hanging up does not stop the charge either. The cost arrives in the final usage chunk and covers the
whole generation, so when you disconnect MiniCrow stops sending but lets the generation finish and charges what it
used.
POST /v1/embeddings
minicrow-embed returns dense and sparse vectors in one call. The sparse half (lexical weights) is what hybrid
retrieval needs, and most embedding APIs do not return it.
{"model":"minicrow-embed","input":["kal meeting hai Pune me","tomorrow there is a meeting"]}
{"object":"list","model":"minicrow-embed","data":[
{"object":"embedding","index":0,"embedding":[…1024 floats…],
"sparse":{"indices":[5,9,…],"values":[0.7,0.2,…]}}],
"usage":{"prompt_tokens":12,"cost":0.0025,"cost_currency":"INR_paise","prompt_tokens_estimated":true}}
⚠️ prompt_tokens_estimated: true is not decoration. Embedding produces no token count of its own, so the
count is MiniCrow's estimate, made with the same per-script estimator MiniCrow uses when it picks a chat mode.
Every other endpoint bills from a measured figure; this one says when it cannot.
|
|
| at most 256 inputs |
capacity is shared with live retrieval; an unbounded batch stalls everybody's search |
| pre-tokenised input is refused |
OpenAI accepts token-id arrays. MiniCrow's tokenizer is not OpenAI's, so those ids mean nothing to it — embedding them would return confident vectors for the wrong text |
a short batch is a 503 embeddings_failed |
the results are positional. Three vectors for four texts would put every vector on the wrong text, with no error anywhere. It is decided as a 502 and answered as a 503 — see Errors |
POST /v1/rerank
Cohere's shape, because OpenAI has no rerank endpoint and every client that speaks rerank speaks this one.
{"model":"minicrow-rerank","query":"Pune meeting kab hai",
"documents":["Delhi ka flight subah 6 baje","Pune me meeting kal 3 baje hai"],"top_n":2}
⚠️ Billed on query + documents, not documents alone. A cross-encoder reads the query once per document, so
counting only the documents would under-count every request with a long query — which is most of them.
Rerank is billed per token at the embedding rate; GET /v1/models lists it.
POST /v1/audio/speech — Pica
{"model":"pica-nano","input":"नमस्ते, आज मौसम बहुत अच्छा है।","voice":"david","mode":"natural"}
Returns audio/wav. The charge rides in headers, because the body is audio and a customer must be able to
read what a synthesis cost without parsing a WAV:
X-Model: pica-nano X-Voice: david X-Mode: natural X-Lane: standard
X-Cost-Paise: 1.1958 X-Cost-Currency: INR_paise
# and, when a request was moved: X-Overflow: pica-small (pica-nano was at capacity) or
# X-Fallback: pica-small (pica-large does not speak that language)
X-Duration-S: 2.05 X-Sample-Rate: 24000
Three tiers, billed per hour of audio
model |
What it is |
Price |
pica-nano |
The eleven built-in voices, your own voices (mcv_…), the standard and expressive lanes, emotions, first_clause and chunk |
₹21 per hour of audio |
pica-small |
Its own Hindi and English voices |
₹32 per hour of audio |
pica-large |
The most natural of the three; every voice speaks Hindi and English |
₹60 per hour of audio |
⚠️ Billed per second of audio produced (since 0.45.0; before that Pica was one tier at ₹32 per 10,000 characters).
The charge is the playing time of the audio Pica delivered to you, measured from the audio itself — ₹21 an hour
is 0.5833 paise a second. A request that fails before any audio is not charged; a stream you hang up on is charged
for the audio made up to that point. At most 5,000 characters per request.
- When
pica-nano is at capacity, a pica-nano request may be spoken by pica-small — only a built-in Hindi voice
with no emotion; it is billed as pica-nano, and the answer says so with X-Overflow: pica-small and the X-Voice that
spoke. Your own voices and emotions are never moved; they wait for pica-nano.
- A language
pica-large does not speak is spoken by pica-small (0.62.3). It speaks Hindi, English,
Marathi, Tamil, Telugu, Gujarati, Bengali and Kannada; a text in Malayalam, Punjabi, Odia or Assamese is served by
pica-small with the voice of the same name where it has one, and the tier's default otherwise. It is billed as
pica-small, and the answer says so: X-Model: pica-small and X-Fallback: pica-small. Nothing is refused and
nothing is read in the wrong language.
- The language comes from the script you send (0.62.3). Before that,
pica-large read every text that was not
Devanagari as English, so Tamil, Telugu, Bengali, Gujarati and Kannada were read with an English tongue.
pica-mini is accepted as pica-small — the tier's name for its first day (0.45.0); the answer's X-Model says
pica-small.
- No
model is pica-nano. ⚠️ One of the eleven voices — or one of yours — is always spoken by pica-nano
and billed as pica-nano, whatever model you name: {"model":"pica-small","voice":"david"} keeps working as it
did before the tiers existed, and X-Model: pica-nano says so.
pica-small and pica-large speak with their own voices — GET /v1/audio/voices?model=pica-small lists them. With
no voice, pica-small uses ira for Indian-script text and zara otherwise, pica-large uses saanvi.
- The
pica-small and pica-large voices were renamed on 0.63.0 (23 September 2026): every pica-small and pica-large voice has a new
id and name, and GET /v1/audio/voices lists only those. The old ids keep working — a request that names one is
spoken by the same voice and answered with the new id in X-Voice. Nothing else about the request changes.
- On
pica-small and pica-large, lane, emotion, chunk and first_clause are 400 option_not_supported,
not ignored. mode is accepted and changes nothing. stream and response_format: "mulaw_8k" work as below.
429 tier_busy — the tier is at capacity right now; retry shortly or use pica-nano. 503 tier_not_deployed
— that tier is not configured on this deployment.
| Field |
|
model |
pica-nano (default), pica-small or pica-large — see the table above |
voice |
a voice of that tier, a built-in or one of yours (mcv_…) — GET /v1/audio/voices lists them all. Omitted on pica-nano, you get david for Devanagari text and robert otherwise |
mode |
normal or natural (default). Delivery style, not the Osprey speed/intelligence ladder. On pica-nano the speech is paced after it is made: 1.30× for Hindi voices, 1.05× for English voices (since 0.61.4 — English at 1.30× ran about 270 words a minute) |
lane |
standard (default, both languages) or expressive (Hindi, and the only one that takes an emotion) |
emotion |
the expressive lane only — neutral · happy · sad · angry · disgust · fear · surprise |
stream |
true streams length-prefixed WAV frames, one per unit (a sentence in natural mode) — see below |
first_clause |
true speaks the first sentence's first clause on its own — a few words before the first frame instead of a whole sentence. Off by default; see First audio below |
chunk |
"clause" speaks every sentence clause by clause. "stream" (stream only, standard lane only) sends each sentence as several frames while it is still being made — see Inside a sentence below. Off by default |
response_format |
"mulaw_8k" on a stream delivers 8 kHz G.711 μ-law frames for a phone line — see below. Any other value is accepted and ignored: the answer is 16-bit PCM WAV at the voice's own rate |
Streaming speech
"stream": true sends each sentence the moment it is synthesised, so the first one can start playing while
the rest are still being made. The body is a sequence of frames, each flushed on its own:
┌──────────────────────────────┬────────────────────────────────────────────┐
│ 8 bytes, unsigned big-endian │ that many bytes: one complete WAV file │
│ length N │ (RIFF header, PCM 16-bit mono) │
└──────────────────────────────┴────────────────────────────────────────────┘
… repeated, then the connection's chunked body ends
- One frame per unit. In
natural mode a unit is a sentence (a sentence ends at . ? ! । or a
newline; a run such as ... or ?! ends it once, at its last mark). In normal mode the whole answer is
one frame, all at once, when the whole text is done. With first_clause or chunk a unit is a clause
(below), in either mode. X-Units says how many frames a whole stream has, before the first one arrives.
- Every frame is a whole, playable WAV. Each carries its own header, so a player can start on frame one without
waiting for a total length. Each ends with the pause that follows its sentence already in the audio — play the
frames back to back, with no gap of your own.
- The sample rate is the same in every frame and is sent up front in
X-Sample-Rate: 24000 on the
standard lane, 44100 on expressive, 24000 on pica-small and pica-large. The headers are the same as for
a whole answer — X-Model, X-Voice, X-Mode, X-Lane — except X-Cost-Paise, X-Cost-Currency, X-Duration-S
and X-Seed, which are not known until the end. The cost of a stream is on its usage row (GET /dashboard usage). Content-Type stays audio/wav; the framing is what this section describes.
- There is no
Content-Length and no end-of-stream frame: the stream is finished when the body ends. A body
that ends inside a frame was cut off.
- The same audio costs the same with or without
stream: per second of audio. Hang up part-way and you pay
for the audio made so far. Hanging up stops the synthesis as soon as your connection closes.
- If synthesis fails before the first frame, the answer is a JSON error (
502 speech_failed, or
503 through the public host), not charged. After the first frame the status is already 200; a stream that
fails later simply has fewer frames than sentences, or ends inside a frame.
- The stream is sent with
Cache-Control: no-cache and X-Accel-Buffering: no, and its connection closes
when it ends (Connection: close) — open a new connection for the next request.
- The frames of a stream are not byte-identical to the one WAV the same request returns without
stream: a
whole answer is levelled and paced as one piece, a stream sentence by sentence.
- With
"chunk": "stream" a unit is several frames, and the stream says so with X-Chunk: stream —
see Inside a sentence below.
First audio: first_clause and chunk
A whole first sentence has to be synthesised before the first frame can leave. Measured on the public host
(2026-09-14, N 60): the gateway's time is about 530 ms + 15 ms per character, so a 50-character Hindi
sentence is ~1.3 s to the first byte. "first_clause": true has Pica speak the first
sentence's first clause — the words up to the first , ; : — or । that is followed by a space, plus
following clauses while the head stays within about six words — as a frame of its own, then the rest of the
sentence, then every later sentence exactly as it would have been. Measured with the same clause cut made
client-side (N 56 pairs): the first audio arrived 437 ms earlier at p50, 1,007 ms at p90.
{"input":"आज शाम चार बजे, राजेश के साथ आपकी कॉल तय है। पॉइंट्स अभी तैयार हो जाएँगे।","voice":"david","stream":true,"first_clause":true}
→ X-Units: 3 and three frames: आज शाम चार बजे, · राजेश के साथ आपकी कॉल तय है। · पॉइंट्स अभी तैयार हो जाएँगे।.
- A first sentence with no clause mark is not cut: you get the plan you would have had, and
X-Units says so.
A colon inside a time (10:30) or a comma inside a number (1,000, 1,25,000) is not a clause mark.
- A clause is never spoken alone when it is under three words or ten characters — it is joined to the next
one (
नमस्ते, आज शाम… is not cut at all). Measured on these voices: a greeting generated on its own was
swallowed or garbled, while the same words inside a longer take were perfect.
- A very short sentence — fewer than 3 words or under 10 characters, such as
नमस्ते! or Okay. — is joined
to the next one (or to the previous one when it is last), for the same reason. This happens on every request,
with or without these fields.
- A full stop inside a number or after an abbreviation is not a sentence end, with or without these fields:
10.5 लाख, डॉ. शर्मा, ए.के., Dr., Rs. 500, No. 7, e.g. stay inside their sentence. After a
number the same letters close the sentence (took 300 ms., paid 500 Rs.), and a.m./p.m. close it
when a capital letter follows. Main St. The… is still read as one sentence.
- The clause frame ends in a short breath (about 150 ms after pacing) rather than a sentence pause.
- Every later sentence is unchanged — same seed, same audio as without
first_clause. Only the first
sentence is a different take, spoken in two pieces; whether the seam is audible is something to judge by ear
on your own voice, which is why this is opt-in.
"chunk": "clause" does the same for every sentence: every clause is its own frame. first_clause is then
redundant.
- Both fields work without
stream too — the one WAV is then joined from the same units — but the point of
them is the first frame of a stream.
- The charge does not change: it is per second of audio.
Inside a sentence: chunk: "stream"
first_clause still waits for a whole clause to be synthesised. "chunk": "stream" does not wait for the unit at
all: Pica decodes the sentence while its speech is still being generated and sends each piece as it is
decoded, so the first frame is the first fraction of a second of the sentence.
{"input":"आपने पिछले हफ्ते टू बीएचके फ्लैट के बारे में पूछा था। पॉइंट्स अभी तैयार हो जाएँगे।","voice":"david","stream":true,"chunk":"stream"}
Content-Type: audio/wav X-Sample-Rate: 24000 X-Units: 2 X-Chunk: stream
X-Units still counts units (sentences, or clauses with first_clause) — it does not count frames. How
many frames a unit becomes is decided while it is decoded and is not known up front (about four for a
50-character Hindi sentence). X-Chunk: stream is how you know frames are not units.
- Every frame is still a whole, playable WAV. Each carries a 16-byte RIFF chunk
rkfc between fmt and data
— little-endian u16 unit (0-based, in order), u16 seq (0-based within the unit), u8 flags (bit 0: the
unit's last frame), 3 reserved bytes. Any WAV reader skips it; you need it only to know where a sentence ends.
- Play the frames back to back, with no gap of your own: the cut between two frames of a sentence is not in
the audio. Only a unit's last frame ends in its pause.
- A whole stream is
X-Units frames with the last bit. A stream that ends after a frame without it ended inside a
sentence.
- The words are the same take as without
chunk: "stream" (the same speech tokens, measured on 60 Hindi and
60 English sentences), decoded in pieces while they are made. It does not sound
identical to the whole-sentence frame: the duration matches within about 1.5 %, the joins between pieces were not
measurable in the delivered audio and transcripts were as accurate, but the timbre differs slightly — judge it by
ear on your voice.
- The level of a sentence is decided before the sentence exists, where a whole-sentence frame is levelled
exactly. Measured: within about 0.5 dB on average and 2 dB at most on a Hindi voice; on an English voice 1.3 dB
quieter on average and up to about 4.7 dB. Consecutive sentences can therefore differ in level by a few dB, and a
voice you created, which we have not calibrated, can be further off.
- Same request, same bytes, whatever the load — as long as the service's streaming configuration is unchanged;
when we retune it, where the frames are cut changes and so do the bytes.
- A sentence too short to be worth cutting (roughly under 36 speech tokens, ~1.4 s of audio) arrives as one frame.
400 chunk_needs_stream without "stream": true; 400 chunk_not_supported on the expressive lane, which
cannot decode inside a sentence; 400 chunk_not_available when Pica cannot stream inside a sentence
right now — send the same request without chunk and you get sentence frames.
- ⚠️ It costs Pica more time per sentence (a sentence is decoded several times while it is
made), so under load other requests wait a little longer for their turn.
- The charge does not change: it is per second of audio.
Telephony wants 8 kHz μ-law. On a stream, "response_format": "mulaw_8k" delivers exactly that:
{"input":"आज शाम चार बजे, राजेश के साथ आपकी कॉल तय है।","voice":"david","stream":true,"first_clause":true,"response_format":"mulaw_8k"}
Content-Type: audio/basic X-Sample-Rate: 8000 X-Encoding: mulaw X-Units: 2
- The framing is the same — an 8-byte big-endian length, then that many bytes — but the payload is raw
G.711 μ-law, one byte per sample, 8,000 per second, no header. Feed a frame to the line 160 bytes (20 ms)
at a time as it comes.
- Each frame is Pica's 24 kHz (or 44.1 kHz on
expressive) WAV frame resampled with a windowed-sinc
low-pass (measured: a 1 kHz tone survives at ≥ 60 dB SNR, anything above the 4 kHz Nyquist is removed by
≥ 60 dB) and μ-law encoded (≈ 38 dB SNR, which is μ-law). Frames are resampled on their own, which is safe:
every frame fades in and ends in its own pause.
- With
"chunk": "stream" the frames of one sentence are resampled as one signal — byte for byte what the
whole sentence in one frame would give — so no cut between them is audible on the line. The last ~8 ms of a
frame wait for the next one, and the sentence's last frame flushes them. The μ-law frames carry no tag (they have
no header to carry it in); X-Units and X-Chunk: stream are sent as for WAV.
- ⚠️ With
"chunk": "stream", each unit ends with an EMPTY frame — an 8-byte length of 0 and nothing after it.
It carries no audio (feeding it to the line plays nothing); it is how you tell where a sentence ends. A whole
stream has exactly X-Units empty frames and ends with one; a stream with fewer ended inside a sentence.
- Without
stream, mulaw_8k is a 400 format_needs_stream: a phone line is fed as the frames come, and
a caller who asked for μ-law must never be handed a WAV at HTTP 200.
- On
pica-small and pica-large the audio is made at 8 kHz directly, so nothing is resampled.
- The charge is the same as for WAV: per second of audio.
⚠️ emotion on the standard lane is a 400, not an ignored field. A flat reading of something you asked
to sound angry, billed as if it worked, is worse than a refusal.
⚠️ An English voice will not read Devanagari — 400 script_mismatch. Measured: it says nonsense, and
nonsense billed as speech is worse than a refusal.
GET /v1/audio/voices
Everything this key may speak with: the eleven built-in voices and the voices this account has made (both
pica-nano), and the voices of pica-small and pica-large. Every voice names its model; ?model=pica-large
lists one tier. A pica-large voice lists "languages":["hi","en"] — its language follows your text. Also the two
delivery modes, the two lanes and which of them has emotions (all three pica-nano only). It never returns a
voice's reference clip or its hash: those name a file on our storage and are the thing an abuser would want.
{"object":"list","built_in_count":11,"own_count":1,
"data":[{"id":"david","name":"David","language":"hi","gender":"male","built_in":true,"model":"pica-nano"},
{"id":"neha","name":"Neha","language":"hi","built_in":true,"model":"pica-small"},
{"id":"archana","name":"Archana","languages":["hi","en"],"gender":"female","built_in":true,"model":"pica-large"},
{"id":"mcv_k4t2…","name":"Anil","language":"hi","kind":"generated","duration_s":14.2,
"created_at":"2026-09-11T04:10:00Z","created_by_key":"…","built_in":false,"model":"pica-nano"}],
"modes":["normal","natural"],"lanes":["standard","expressive"],"own_voice_lanes":["standard"]}
POST /v1/audio/voices — make a voice
Two ways in, and they are not the same thing.
| You send |
What happens |
Use it for |
{"description": "warm, unhurried, mid-forties"} as JSON |
a new voice is generated from your description. It copies nobody — it invents a speaker who has never existed |
a brand voice, a narrator, anything where you want a person who is not a person |
a reference recording as multipart/form-data |
your clip becomes the voice, cloned |
a voice you own and want to keep using |
# generate
curl https://api.minicrow.com/v1/audio/voices -H "Authorization: Bearer mc_..." \
-H "Content-Type: application/json" \
-d '{"name":"Anil","description":"warm, unhurried, mid-forties","language":"hi"}'
# clone
curl https://api.minicrow.com/v1/audio/voices -H "Authorization: Bearer mc_..." \
-F file=@reference.wav -F name=Anil -F language=hi -F attest=true
Answers 201 with the voice. Its id is then just a voice on POST /v1/audio/speech.
| Field |
|
description |
JSON only. A sentence or two, at most 600 characters. Describe the voice, not the words |
file |
multipart only. 10 to 20 seconds of clean speech works best; 3 to 30 is accepted. WAV, MP3, FLAC, OGG, AIFF or CAF — m4a, aac and webm are not read |
attest |
multipart only, and required: attest=true |
name |
what you want to call it. Yours; we never show it to anyone else |
language |
hi (default) or en. It chooses the pronunciation lane. hi reads both scripts; en refuses Devanagari, exactly as a built-in English voice does |
What cloning a voice means here, stated plainly
⚠️ attest=true means you are saying: I hold the rights to this recording and have the speaker's permission
to clone their voice. Sending it is a claim you are making, on the record, with a timestamp and the address
it came from.
⚠️ MiniCrow does not verify that claim, and cannot. Nobody can tell from a clip whether the person in it
agreed. There is no consent check behind this endpoint, and calling the attestation a safeguard would be
selling you one that does not exist. What is real is this:
- every attempt — accepted and refused — is logged against the API key that made it;
- the reference recording is kept, so a complaint about a voice can be answered with the audio it was made
from;
- creation is rate-limited per key (10 an hour, 30 a day by default), and an account holds at most 100
voices at a time.
Whose voice it is
⚠️ A voice belongs to the account that made it and is never visible to any other customer. Another
customer cannot list it, speak with it, or delete it — and an id that is not yours answers 404 unknown_voice,
worded identically to a typo, because a 403 would confirm the id exists. The id itself is 128 random bits, so
it cannot be guessed or walked.
⚠️ Every key on your account can use every voice on your account. The key that created a voice is recorded
and is what the rate limit and any complaint are measured against — but rotating a key does not lose you your
voices.
⚠️ A voice you made speaks on the standard lane only. The expressive lane has its own fixed cast, so a
custom voice with lane: expressive is a 400 lane_not_supported rather than a stranger's voice billed as
yours.
DELETE /v1/audio/voices/{id}
{"object":"voice.deleted","id":"mcv_k4t2…","deleted":true}
The voice stops working immediately — a synthesis naming it afterwards is 400 unknown_voice.
⚠️ The reference recording is retained after deletion. That is the deliberate cost of being able to answer
a rights complaint about a voice weeks after somebody deleted it. Deleting removes the voice from service; it
does not erase the recording it was made from. If you need the recording itself removed, ask us.
Deleting a built-in voice is 400 voice_not_deletable: the eleven belong to MiniCrow and are available to
every account.
POST /v1/audio/transcriptions — Lark
multipart/form-data, OpenAI's shape.
curl https://api.minicrow.com/v1/audio/transcriptions \
-H "Authorization: Bearer mc_..." \
-F file=@call.wav -F model=lark-mini
{"text":"Kal shaam 5:00 baje Pune me meeting hai.","model":"lark-mini",
"x_minicrow":{"tier":"lark-mini","cost_known":true},
"usage":{"seconds":2.05,"cost":0.3928,"cost_currency":"INR_paise",
"duration_estimated":false,"audio_format":"wav"}}
⚠️ Billed per second of audio, and the seconds come from your recording — never from a field you send.
A client-supplied duration is a number anybody can set to 1. MiniCrow reads the container's own header: exact
for WAV and OGG/Opus, estimated for MP3/M4A/WebM, and duration_estimated tells you which.
⚠️ Code-mixed Hindi comes back in Latin script, not Devanagari. Left alone, a transcript of code-mixed Hindi
tends to come back in Devanagari; Lark writes it as the speaker said it.
⚠️ prompt is background about the recording; it cannot change the task. Use it for a name spelling or a domain
hint. Whatever it says, what comes back is a transcript of your audio.
| Tier |
|
lark-nano |
coming soon — the economy lane (₹10/hour, at most 5 minutes of audio per request). Until it opens it answers 503 tier_not_deployed; use lark-mini |
lark-mini (default) |
the balanced lane — ₹20/hour today, whichever language you send. For Hindi and English, which lane hears you depends on your recording's sample rate and on upload vs live (below) |
lark-large |
the premium lane — ₹30/hour, or ₹45/hour for an Indian language |
lark-large:max |
the top lane for both language groups — ₹60/hour |
⚠️ One price per tier, not per lane — with one exception. lark-mini runs two lanes and charges the
same for both: which one hears you is our decision and our cost, not a line on your bill. lark-large is the
exception and says so above, because its Indian-language lane genuinely costs more to run.
⚠️ The price may come to differ by sample rate and by streaming. lark-mini now sends telephony-rate Hindi and
English to its dearer lane (below), and Lark Live's lanes are separate from the upload endpoint's. Today both
are billed at the one ₹20/hour; when a rate changes it changes here, in the GET /v1/models price, and in the
changelog — never silently on an invoice. Read the price for the sample rate and the endpoint you actually use.
The language you declare can change the lane
-F model=lark-mini -F language=hi # Hindi: telephone audio (up to 16 kHz) -> lane `indic`;
# audio above 16 kHz -> lane `default`
-F model=lark-mini -F language=en # English: the same rule as Hindi
-F model=lark-mini -F language=mr # other Indian languages -> the Indic lane at every rate — same ₹20/hour
-F model=lark-large -F language=mr # an Indian language, Hindi included -> ₹45/hour
-F model=lark-large -F language=en # anything else -> ₹30/hour
# nothing declared -> the `default` lane
⚠️ On lark-mini, Hindi's and English's lane follows your recording's sample rate, and that is measured,
not tidy. Telephone audio — 8 kHz and 16 kHz, which is what a call recording is — goes to lane indic; audio
above 16,000 Hz (a microphone recording at 44.1 or 48 kHz) goes to lane default. A voice note can go either
way: an OGG/Opus note is judged by the input rate its header declares, and many phone apps (WhatsApp among them)
record voice notes at 16 kHz, so those take lane indic; check audio.sample_rate_hz in the response. On
100 Hindi clips with human-written references pushed through a telephone channel, the indic lane was a full
point better in word error rate and the default one returned nothing at all on four of the hundred; on 100
English clips through the same channel the gap was far wider. At 48 kHz the two had measured level on Hindi.
The line is 16,000 Hz: at or below it is telephony. It is a rule we can move without a release, so the numbers
here are today's. lark-large has no such rule — Hindi is an Indian language there at every rate, and English
takes its default lane.
⚠️ indic is the name of a lane, not a claim about your language. On lark-mini it is the lane that hears
telephone audio best, and English telephone audio is sent to it for that reason. x_minicrow.lane: "indic" on
an English call is expected.
⚠️ The rate is read from your file's own header, never decoded, and a rate we cannot read takes the
telephony lane. WAV (fmt ), OGG/Opus (the encoder's input rate — Opus itself always runs at 48 kHz, so
the header's input rate is what your recorder was fed) and MP3 (the frame header) are read. M4A and WebM
are not opened that far: Hindi and English in those containers take lane indic — the language's base rule — and
the response carries no audio.sample_rate_hz, so you can tell "not read" from "read and below the line".
Send WAV or OGG if you want the rate to decide.
⚠️ Live is telephony by construction. The live socket takes 8 and 16 kHz only, so live Hindi always uses
the Indic lane's final pass; the lane is hindi and its price is in the live section.
⚠️ The same file can therefore cost differently depending on how you send it. Today every one of these is
₹20/hour; the price is set per lane and per endpoint, so it may come to differ by sample rate and by streaming.
Check GET /v1/models and this page for the endpoint and the rate you actually use.
⚠️ You declare it; we do not guess. Language cannot be detected before transcription, and a wrong guess
sends Marathi to a lane that has never heard it. A request with no language takes the default lane —
which is also the cheaper of the two to get wrong on lark-large.
The response says which one served and why, so you can tell without asking:
"x_minicrow": {"tier":"lark-mini","lane":"indic","lane_reason":"language",
"audio":{"container":"wav","channels":1,"sample_rate_hz":8000}}
lane_reason |
|
mode |
you named a mode (lark-large:max) |
language |
the language you declared chose the lane |
default |
no rule applied; the tier's default lane |
sample_rate |
the language's rule has a second lane above a sample rate, and your recording was above it — today, Hindi and English on lark-mini above 16 kHz |
audio.sample_rate_hz is the rate that was read; it is absent when the container did not say, and then the
language's base lane served.
Indian languages: hi mr gu bn ta te kn ml pa or as ur ne sa kok mai sd ks doi mni sat brx bho raj
(on lark-mini, hi — and en — by sample rate as above).
The alphabet is yours to choose
# nothing sent -> the language as it is WRITTEN. The default.
-F script=native # Devanagari for Hindi and Marathi, each language in its own script
-F script=latin # romanised — "kal shaam ko meeting hai"
-F script=auto # whatever Lark writes, with no alphabet asked for either way
devanagari, original and source are synonyms for native; roman and romanised for latin.
Anything else is a 400 unknown_script rather than a silent fall back — a caller who typed devnagari and
received romanised text could not tell that from an API that ignored them.
⚠️ native keeps your code-mixing, it does not translate it into Devanagari. An Indian business call is
often half English: the speaker says meeting, link, online, payment inside a Marathi sentence.
Those stay in English, in Latin letters, exactly as spoken — you get उद्या meeting आहे का?, never
उद्या मीटिंग आहे का?. Numbers a speaker says as numbers come back as digits. Nothing is translated.
The response tells you which alphabet you asked for and what happened:
"x_minicrow": {"script": "native", "script_honoured": true, "script_repaired": false}
⚠️ A language written in the Latin alphabet is already in its own script. For en, fr, es and the other
Latin-written languages, native means Latin: an English transcript is script_honoured: true and is never
rewritten. With no language declared, an all-Latin transcript is not judged (script_honoured absent) —
English and romanised Hindi look the same, and we will not guess. So a romanised Hindi transcript is only rewritten
into Devanagari when you declare language=hi (or another Devanagari-written language).
⚠️ script_repaired: true means we had to fix it, and you were not charged for that. About one Indian-
language recording in seven is first transcribed in the wrong alphabet — the words correct, the letters
not — and retrying reproduces it exactly. Rather than hand you a writing system you did not
ask for, the transcript is rewritten into the one you did. It is a script conversion and nothing else: no
word is translated, corrected, added or removed. Both ways it can go wrong are repaired: English words
written in Devanagari go back to English letters, and Hindi or Marathi written in English letters
(make sure karna ki) goes into Devanagari (make sure करना कि). A repair that would change the words is
refused; the transcript then comes back as it was, with script_honoured: false.
Who spoke — speaker labels
-F diarize=true
-F speakers=2 # optional: how many people are on the recording, if you know
Adds ₹3.50 per hour of audio on top of the tier's own rate, and returns the turns beside the transcript:
{
"text": "Hello sir, mi Pooja bolte Samarth Sky project madhun. Ha bola, kay aahe sanga...",
"segments": [
{"start": 0.0, "end": 4.2, "speaker": 1},
{"start": 4.5, "end": 7.1, "speaker": 2}
],
"x_minicrow": {"speakers": 2}
}
⚠️ The transcript is unchanged. Labels do not get interleaved into text — that would break every caller
already parsing it, and would throw away the timings, which are the half a CRM actually needs. Join them
yourself on the times.
⚠️ Tell us speakers when you know it. A two-party phone call is two people; leaving Lark to work
that out costs accuracy you could have had for free.
⚠️ If the speaker pass fails, you still get the transcript and you are not charged for the labels. The
words were already correct and already earned; segments comes back empty and the ₹3.50 is not applied.
How good is it? Measured on constructed two-speaker telephone audio with exact ground truth: 100% of
frames attributed to the right speaker, three recordings, and roughly 25× faster than real time. That is a
ceiling, not a promise — those were two clearly different voices with little overlap. Real calls have similar
voices, crosstalk and a worse line; published figures for speaker labelling of this kind land nearer 85–90%, and we have
not yet measured it on a hand-labelled real call.
The single biggest thing you can do is record each party on their own channel. Then there is nothing to
infer — the channel is the speaker — and it costs nothing.
Two channels, two speakers
An upload may hold one speaker or many, and we do not assume a two-channel file is two speakers. A
two-channel WAV (16-bit PCM, μ-law or A-law) is transcribed channel by channel only when the file shows it:
each channel carries speech of its own — energy above 1 kHz, a spectrum that keeps moving, loudness that rises
and falls at the rate of syllables — and the two channels are not the same signal (correlation below 0.9 at every
alignment within ±50 ms, so a copy of one channel written a few milliseconds late counts as the same signal). A
silent channel, a steady tone or line hum on one side, or the same mix written to both channels fails, and the
file is transcribed as one stream exactly as a mono file is.
⚠️ What the check cannot tell apart. It measures whether each channel sounds like speech, not whose speech
it is, so: music on one channel (hold music with a beat) can pass at 8 and 16 kHz, and steady background hiss
can pass too — that channel is then transcribed on its own and comes back empty or near-empty. And a panned
mix — both voices on both channels at different levels, e.g. one voice at full level and the other at half —
correlates below 0.9 and splits, so each channel's transcript repeats most of the conversation. If your recorder
pans rather than separates, send a mono file. Compressed two-channel files (OGG, MP3, M4A, WebM) are
not opened that far and always take the one-stream path.
When it splits, both channels go to the same lane, and the response labels them:
{
"text": "channel_1: Hello sir, mi Pooja bolte...\nchannel_2: Ha bola, kay aahe...",
"segments": [
{"start": 0.0, "end": 64.2, "channel": 1, "speaker": "channel_1", "text": "Hello sir, mi Pooja bolte...",
"script_honoured": true, "script_repaired": false},
{"start": 0.0, "end": 64.2, "channel": 2, "speaker": "channel_2", "text": "Ha bola, kay aahe...",
"script_honoured": true, "script_repaired": false}
],
"x_minicrow": {"channel_split": {"split": true, "correlation": 0.12,
"rule": "each channel was transcribed on its own on the same lane; audio seconds are billed once"}}
}
⚠️ Channels count from 1: channel 1 is the left channel, channel 2 the right, in text, segments and
channel_split's reasons alike.
⚠️ The alphabet is judged channel by channel. Each segment carries its own script_honoured and
script_repaired; x_minicrow.script_honoured is true only when every channel that can be judged is, and
script_repaired is true when any channel was repaired.
⚠️ Billed once, by the length of the recording. A 64-second two-channel call is 64 billed seconds, not 128.
⚠️ Each segment spans the whole call. A channel's words are not cut into timed turns; speaker is a
channel label, a string, where the speaker pass (below) returns a number.
⚠️ diarize=true on a file that splits is skipped, not charged, and not refused. The channels already
separated the speakers; x_minicrow.diarization says {"skipped": true, …}. On a file that does not split,
diarize works exactly as described above.
When a two-channel WAV does not split, x_minicrow.channel_split says why:
{"split": false, "correlation": 0.998, "reason": "both channels carry the same audio; transcribed as one stream"}.
A mono file has no channel_split at all.
⚠️ Live is one speaker per stream. The live socket takes one channel; for a two-party call open one session
per leg.
The transcript is of your recording, never one of Lark's own lines
On a recording with little or nothing in it, Lark can answer with one of its own reference lines instead of a
transcript. MiniCrow checks every transcript against those lines; a transcript that is one — at least 80 % of
the line's words in a row, making up at least 80 % of the transcript — is retried once, and if the retry does the
same, the request fails with transcription_failed (HTTP 503 at our edge) and nothing is charged — the line
never reaches you. A transcript that merely contains such a line's words, such as a real call that opens with the
same line and carries on, is returned as normal. ⚠️ A recording whose entire content is one of those lines, word
for word, is indistinguishable from a copy and is refused.
Declaring the language improves accuracy
0.35.3 (pending). Lark's default handling is tuned for Hindi/English code-mix and for Marathi. That helps Hindi
and Marathi, and it hurts other languages: on other Indian languages Lark sometimes wrote the transcript in
Devanagari, or in Marathi, or romanised it. So when you declare one of these twelve languages, Lark transcribes it
as that language alone, word for word and never translated:
language |
what you get (with script=native or no script) |
ta te bn gu kn ml |
the language in its own script (Tamil in Tamil script, and so on); English words the speaker says stay in English |
en es de fr it ar |
the language as it was spoken, not translated |
Measured on 100 FLEURS clips per language through an 8 kHz telephone line, on the indic lane: character error
fell by 16.4 points on Kannada, 8.8 on Gujarati, 5.8 on Bengali, 5.6 on Tamil, 4.7 on Malayalam and 3.5 on Telugu,
and word error fell by 1.8 on Arabic. On all of these, no clip came back in the wrong alphabet (38 of 600
Indian-language clips did before). English, Spanish, German, French and Italian scored the same either way.
hi, mr, no language, and every other language keep the default handling, unchanged. For Hindi and
Marathi, the per-language handling wrote English words in Devanagari (फ्लॅट), and your code-mixed calls need
them in Latin. That has to be checked on real calls before those two change.
script=latin on Tamil, Telugu, Bengali, Gujarati, Kannada, Malayalam or Arabic asks for the Latin alphabet
(romanised) instead of the language's own script. script=auto names no script at all. On English, Spanish,
German, French and Italian, latin is the same as native. Neither variant has been measured.
domain, vocabulary, abbreviations and prompt work the same way with these languages. For a language whose
script is named, the script you asked for always has the last word.
- The copy check above and the alphabet repair are the same for every language.
Language packs
0.37.0 (pending). Switched on lane by lane; until your lane is switched, everything above applies unchanged.
A language pack tunes Lark to one declared language alone — its script, how English words are written in it, its
numerals and punctuation — and to nothing about any business. Packs exist for 14 languages:
language |
how English words the speaker says are written (script=native or no script) |
mr hi bn ml |
in the language's own script, as a newspaper in that language prints them (बस, ऑफिस) |
ta te gu kn |
in English letters; a borrowed word that is part of everyday speech stays in the language's script |
en es de fr it ar |
the language's own spelling and punctuation, numbers as digits; Arabic without added vowel marks |
⚠️ For mr, hi, bn and ml this changes what your transcript looks like. Without a pack an English word
such as "flat" comes back in Latin letters inside a Devanagari sentence; with the pack it comes back in
Devanagari. If you store or search code-mixed text with English words in Latin letters, ask for script=latin
(which keeps today's behaviour) or tell us before your lane is switched.
- Only with
script=native or no script. script=latin on an Indian language, script=auto, no language,
and every language without a pack keep the behaviour above, unchanged. On en es de fr it, latin is the
same as native.
- Measured on 50 held-out read-speech clips per language through an 8 kHz telephone line: without the packs,
Lark trailed a leading Indian speech-to-text service by 1.75 words per hundred pooled over nine languages; with
the packs it was level with it in every one of those languages. ⚠️ That measurement also gave each clip a draft
from Lark Live, which this endpoint does not do. Without the draft, the pack set that
ta te gu kn en es de fr it ar use was ahead of no pack by 0.43 words per hundred, pooled over 700
clips in all 14 languages (not held out, and only just outside the margin); the set mr hi bn ml use has not been
measured without the draft.
- The copy check compares your transcript with the pack's own reference lines, and only those. The alphabet repair
is not applied:
script_honoured is true when the transcript carries letters of the language's own script (for
en es de fr it, when it carries no other script).
domain, vocabulary, abbreviations and prompt work with a pack as they do without one, and never override
it.
- On Lark-V, the speech half uses the pack when the Lark-V tier is switched, independently of the audio tiers.
Tell us about your own recording
Lark's own handling is the same for every customer: it knows your language and script and nothing about your
business. What your recordings are about comes only from you, in these optional fields:
-F domain="<who is talking to whom, about what — one line>"
-F vocabulary="<names, places, products and codes that will be said, as you write them>"
-F abbreviations="<SHORT = long form>, <SHORT = long form>"
-F prompt="<anything else, one or two sentences>"
For example, three different businesses:
|
domain |
vocabulary |
abbreviations |
| a clinic |
appointment calls to a family clinic |
Dr. Meera Iyer, Dr. Arjun Rao, CBC, lipid profile |
OPD = outpatient department |
| a courier |
delivery support calls for a courier company |
AWB, Andheri hub, Swift Express |
COD = cash on delivery, RTO = return to origin |
| a lender |
loan servicing calls |
NACH, foreclosure, Kotak |
EMI = equated monthly instalment |
The rule book
- Send nothing you do not have. Lark is built to work with no fields at all.
domain is one plain line: who is speaking and about what. A description, not a list of rules.
vocabulary is for this recording. Build it per call from what you already know: the customer's name, the
product, the branch — names that will be said. A person's name that is never spoken can be written into a
greeting nobody said ("मैं बोल रहा हूं"), so leave out an agent's name unless the agent says it on the
call — or send verify_vocabulary=true (below). Words a general listener would spell correctly anyway do not belong.
- Spell each entry the way you want it written. The entry is copied letter for letter when it is heard.
abbreviations say how a short form is written, SHORT = long form. The short form is what is written.
- Do not use
prompt for language or alphabet. language and script decide those. prompt is
background about the recording; it cannot change the task.
- Never paste your agent's script, your FAQ or a previous transcript. Hints are text Lark reads next to
your audio; on a turn with no clear speech, a list of names has come back as the transcript.
- Keep it short. Every field is processed with every request.
These are separate fields rather than prose in prompt because their handling is measured: they are used the
same way on every request, where a caller rewording their own prose gets a different result each
time. prompt still works, as background.
⚠️ All four fields are optional, and more is not better. Lark is built to work with none
of them. Measured on read speech in 14 languages: eight vocabulary words and a one-line domain moved accuracy by
less than the measurement's margin in most languages; on real calls, hints helped only where the listed names were
actually spoken. A thick prompt can lower accuracy: a 600-word brief lowered accuracy by 3 to 6 points on 102 real
calls on one Lark lane (on another the same brief raised accuracy by about 4 points on those
calls, so the effect depends on the lane and on how well the brief matches the audio), and a 150-entry glossary unrelated to
the audio made errors lean worse in Hindi and Tamil. Every hint is read with every request, so a long one also costs more to process. Send the names you expect in
this recording; leave the rest out.
⚠️ A vocabulary is a spelling guide, not a script, and it is enforced as one. Measured: a property call
given a cardiology glossary wrote zero of those medical terms into the transcript across 30 runs, while
the correct glossary raised the density of correctly-spelled domain names by 47%. Send the names this
recording actually contains — a list longer than the transcript stops being a hint, and on clear English
audio a long glossary measurably costs accuracy.
Checking the vocabulary: verify_vocabulary
A vocabulary name can take over a stretch of audio that sounds a little like it — an unclear greeting, a similar
name. Send -F verify_vocabulary=true and every place the transcript writes one of your entries is checked:
- If the transcript writes none of your entries, nothing more happens and nothing more is charged.
- Otherwise the recording is transcribed a second time without your
vocabulary — domain, abbreviations,
prompt, language and script are kept — and the two transcripts are lined up word by word, by sound.
- An entry the second transcript heard at the same place, in any spelling (
नासिक for Nashik), stays spelled as you
listed it. An entry where the second transcript heard something else is replaced by what it heard, together
with the words around it the two disagree on: मैं Rohan Mehta बोल रहा हूं becomes मैं देख रहा हूं when that is
what was said.
- Anything in doubt stays as written: an entry where the second transcript heard nothing at that place, a stretch
the two cannot be lined up on, an acronym (
BHK), a name in a script other than Latin or Devanagari.
"x_minicrow": {
"vocabulary_check": {"ran": true, "confirmed": ["Nashik"], "removed": ["Rohan Mehta"], "unverified": []}
},
"usage": {"seconds": 47.1, "cost": 52.22, "vocabulary_check_cost": 26.11, "cost_currency": "INR_paise"}
Price: the second transcription is billed as one more transcription of the recording — its seconds, once, at the
same tier's rate — only when it ran and answered. usage.vocabulary_check_cost is that part and is included in
usage.cost. On a two-channel call only the channels that wrote an entry are transcribed again, and the recording's
seconds are billed once more. When the check does not run, vocabulary_check is {"ran": false, "reason": …} (no
vocabulary sent; no entry written; the second transcription failed; a lane that does not use a vocabulary, such as
lark-large:max) and nothing extra is charged. On a split call each channel's segment carries its own
vocabulary_check.
Measured: on 50 real telephone calls, each with the names spoken on it as the vocabulary, the check kept 129 of 132
spoken names written, left 3 as written, and took out none; the one name written over other words in the test set
was replaced by what was said. It does not catch a listed name that sounds like the one said — Mehta written for a
spoken मेहरा — because at that closeness it cannot tell a wrong name from a spelling. Off unless you ask for it.
⚠️ A recording with no speech is not transcribed, and costs nothing. A WAV that certainly carries no speech —
silence, a steady tone or hum, flat line noise, a busy or ringing tone — is not sent to a model: it comes back with
"text": "", x_minicrow.no_speech: true and usage.cost: 0. It is judged on the audio itself, for WAV files of
3 seconds or more; anything in doubt is transcribed as usual. On a two-channel call transcribed channel by channel,
each channel is judged on its own: a channel that certainly holds no speech comes back as an empty segment with
"no_speech": true and is not sent to a model, and when no channel holds speech the whole call is answered as above.
Any other two-channel file is answered this way only when every channel is without speech. A tone
that switches on and off — as some lines do while the other side talks — can still pass as speech and be
transcribed: measured on 634 recordings, no rule we tried separated such a tone from real speech without also
dropping real speech. x_minicrow.channel_split only says whether each channel looks like one speaker; it is not a
no-speech verdict.
⚠️ A transcript that repeats itself is never returned. If the model falls into a loop — one sentence written
over and over — the recording is transcribed once more without your domain, vocabulary, abbreviations and
prompt, and that transcript is returned if it does not loop; otherwise the repeated words are kept once. Either
way x_minicrow.repetition says so ("retried_without_hints" or "collapsed"); it is absent on every other
response. It only fires on a run no speaker makes — one unit of words repeated back to back at least four times and
at least twenty words long — and the recording is billed once.
Oversize is refused, never truncated: domain 120 characters, vocabulary 2,000 characters or 150 entries,
abbreviations 1,200 characters or 100 entries, prompt 2,000 characters.
On every other tier language is accepted and does not move the lane, because there is only one lane to
move to. The lane in the response always says which one served.
What this endpoint refuses
At most 25 MB per request, matching OpenAI's limit. Beyond that:
| Status |
Code |
When |
| 400 |
file_required |
no file part in the form |
| 400 |
unsupported_audio |
the container is not one MiniCrow can measure. Send WAV, OGG/Opus, MP3, M4A or WebM — the charge is per second and the seconds come from the file's own header, so a format we cannot read is one we cannot bill. FLAC and raw PCM land here |
| 400 |
empty_audio |
the file is a valid recording of zero seconds. It is refused rather than transcribed: a model handed silence will confidently write words nobody said, and a free transcription is indistinguishable from a working one until somebody reads an invoice |
| 400 |
file_unreadable |
the part is empty or could not be read |
| 400 |
unknown_model / unknown_mode |
no such tier, or a mode that tier does not offer |
| 413 |
file_too_large |
over 25 MB; split the recording |
| 413 |
audio_too_long |
lark-nano only — at most 5 minutes of audio in one request. 25 MB is a size limit, not a duration one: 25 MB of Opus is nearly an hour. Split the recording, or send it to lark-mini, which has no such limit |
| 503 |
tier_not_deployed |
the tier, or the mode, is not serving on this deployment. The message names what to use instead |
GET /v1/audio/transcriptions/live — Lark Live
Preview — enabled per account. Lark Live is open to named accounts while it is measured on real traffic.
On any other account every connection is refused 503 tier_not_deployed. Ask us to enable yours.
A WebSocket. You stream a phone call's audio as it happens; MiniCrow finds where each speaker's turn ends and
answers every turn twice — a fast draft the moment the turn ends, then a checked final once the turn
has been listened to again. On intl languages you also get words while the turn is still being spoken.
wss://api.minicrow.com/v1/audio/transcriptions/live?model=lark-mini&language=mr&script=latin&encoding=mulaw&sample_rate=8000
Authorization: Bearer mc_...
⚠️ The key goes in the Authorization header, and nowhere else. Not in the URL — a query string is written
into proxy and server logs — and not in a first message. A browser's WebSocket cannot set a header, so connect
from your server, never from a web page: a key in a page is a key anybody can read.
Python
# pip install "websockets>=14"
import asyncio, json, os, wave
from urllib.parse import urlencode
import websockets
params = urlencode({
"model": "lark-mini", "language": "mr", "script": "latin",
"encoding": "linear16", "sample_rate": 16000,
"domain": "delivery support calls for a courier company",
"vocabulary": "AWB, Andheri hub, Swift Express",
})
URL = f"wss://api.minicrow.com/v1/audio/transcriptions/live?{params}"
HEADERS = {"Authorization": f"Bearer {os.environ['MINICROW_API_KEY']}"}
async def send_audio(ws, path):
with wave.open(path, "rb") as w: # 16 kHz, mono, 16-bit PCM
assert (w.getframerate(), w.getnchannels(), w.getsampwidth()) == (16000, 1, 2)
while chunk := w.readframes(1600): # 100 ms of audio
await ws.send(chunk) # bytes are sent as a binary frame
await asyncio.sleep(0.1) # real-time pace, as a live call arrives
await ws.send(json.dumps({"type": "end"}))
async def main():
try:
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
sender = asyncio.create_task(send_audio(ws, "call-16k.wav"))
async for message in ws: # ends when the server closes with 1000
event = json.loads(message)
if event["type"] == "transcript.partial":
print(f" ({event['stage']}) {event['text']}")
elif event["type"] == "transcript.final":
print(f"[{event['start']:.2f}-{event['end']:.2f}] {event['text']}")
elif event["type"] == "session.ended":
print(f"billed {event['billed_seconds']} s, {event['cost']} paise")
elif event["type"] == "error":
print("error:", event["code"], event["message"])
await sender
except websockets.exceptions.InvalidStatus as refused: # refused before the socket opened
print(refused.response.status_code, refused.response.body.decode())
asyncio.run(main())
JavaScript (Node)
// npm install ws — save as live.mjs, run with node live.mjs
import fs from "node:fs";
import WebSocket from "ws";
const params = new URLSearchParams({
model: "lark-mini", language: "hi", script: "latin", encoding: "mulaw", sample_rate: "8000",
});
const ws = new WebSocket(`wss://api.minicrow.com/v1/audio/transcriptions/live?${params}`, {
headers: { Authorization: `Bearer ${process.env.MINICROW_API_KEY}` },
});
// Refused before the socket opened: the body is the usual JSON error.
ws.on("unexpected-response", (req, res) => {
let body = "";
res.on("data", (chunk) => (body += chunk));
res.on("end", () => console.error(res.statusCode, body));
});
ws.on("open", async () => {
const audio = fs.readFileSync("call.ulaw"); // raw 8 kHz μ-law, no header
for (let i = 0; i < audio.length; i += 800) { // 800 bytes = 100 ms
ws.send(audio.subarray(i, i + 800)); // a Buffer is sent as a binary frame
await new Promise((resolve) => setTimeout(resolve, 100));
}
ws.send(JSON.stringify({ type: "end" }));
});
ws.on("message", (data) => {
const event = JSON.parse(data.toString());
if (event.type === "transcript.final") console.log(`[turn ${event.turn}] ${event.text}`);
if (event.type === "session.ended") console.log(`billed ${event.billed_seconds} s, ${event.cost} paise`);
if (event.type === "error") console.error("error:", event.code, event.message);
});
ws.on("close", (code, reason) => console.log("closed", code, reason.toString()));
Query
Everything is decided before the socket opens. A wrong parameter is an ordinary JSON error on the upgrade
request, not a socket that opens and then closes.
| Parameter |
Values |
|
model |
lark-mini (default) |
the only tier with a live lane |
language |
required. mr bn te ta gu kn or ml pa as · hi · en es fr de pt it ru ar ja ko |
a region is accepted and ignored — mr-IN is mr. Any other language is refused, never guessed |
script |
latin or native — required for every Indian language, Hindi included |
the same synonyms as the batch endpoint (roman, devanagari, …). auto is refused on the live lane. Ignored for en es fr de pt it ru ar ja ko, which are written in their usual script |
encoding |
linear16 (default) · mulaw · alaw |
linear16 is signed 16-bit little-endian mono PCM. μ-law and A-law are G.711, as a telephone line carries them |
sample_rate |
8000 · 16000 |
mulaw and alaw are 8000 only. Omitted, it is 16000 for linear16 and 8000 for mulaw/alaw |
endpointing_ms |
300–1000, default 400 |
how much silence ends a turn. Lower answers sooner and splits more sentences in two |
vad |
server (default) · client |
client: you decide where a turn ends by sending {"type":"turn_end"} |
domain |
at most 120 characters |
what the calls are about, one plain line — delivery support calls for a courier company |
vocabulary |
at most 40 terms, 2,000 characters, comma-separated |
names as they are written. Only written into a transcript when they are said. 0.41.0: also boosted inside the draft on hi, mr and en — the languages the boost is measured on (see Hearing vocabulary under Osprey Live: names not words; for hi/mr add the Devanagari spelling as its own term). 0.43.0: hi and mr drafts are locked to Devanagari, vocabulary or not — English words in English letters, digits and punctuation stay |
context |
repeatable, at most 3 values of 500 characters each |
the last finals of an earlier session, when you reconnect mid-call |
The response says which lane serves you: indic for mr bn te ta gu kn or ml pa as, hindi for hi, intl
for the rest. It is decided by language alone — as on the batch endpoint, you declare it and we do not guess.
⚠️ domain, vocabulary and context are optional. They are used on every turn, so a long vocabulary
costs more on every turn and can add latency to every final; eight words and a one-line domain measured within
40 ms. Send names that will actually be said. A list of names can also be read back as a final on a turn with no
clear speech. The rule book under Tell us about your own recording applies here too; on a live call, build the
vocabulary per call (the customer's name, the product, the branch) rather than one list for every call.
0.37.0 (pending), language packs on live. When your live lane is switched, a session with script=native in
hi ta bn gu ml (English words written in the language's own script) or mr te kn (English words in English
letters) — and every session in en es de fr it — gets the same kind of domain-free language pack as the batch
endpoint (see Language packs above). script=latin, and every other language,
Arabic included, keep today's live behaviour. A session keeps the pack it started with, or none, to its end. The
earlier finals it carries are the last three that were not empty.
⚠️ Speaker labels are not offered live yet. diarize or speakers on this endpoint is
400 diarize_not_available_live. Record each party on its own channel and open one session per channel (each
is billed) — then the session is the speaker.
What you send
| Message |
|
| binary frame |
audio in the encoding and sample_rate you declared, 20 ms to 1 s per frame. A frame over 64 KiB closes the session with 1009 |
{"type":"turn_end"} |
vad=client only: the turn in progress ends here |
{"type":"end"} |
no more audio. Every open turn is finished, then session.ended, then close 1000 |
{"type":"ping"} |
answered with {"type":"pong"} |
Any other text message is answered with a non-fatal error bad_message and the session carries on.
You may send a recording faster than real time. While 16 turns are still waiting for their finals we stop reading
your socket until they are sent, so your writes slow down instead of piling up. One session carries at most 3 hours
of audio, however fast it arrives.
⚠️ Send end, do not just hang up. end waits for the last turn's final. A dropped socket is still billed for
every second it delivered, and the finals still in flight are lost.
What you receive
One JSON object per message, in this order for every turn: speech.started → transcript.partial →
transcript.final.
{"type":"session.started","session_id":"6f1c…","model":"lark-mini","lane":"indic","language":"mr",
"script":"latin","encoding":"mulaw","sample_rate":8000,"endpointing_ms":400,"vad":"server"}
{"type":"speech.started","turn":0,"start":0.42}
{"type":"transcript.partial","turn":0,"stage":"draft","text":"हो सर फ्लॅट चा बुकिंग अमाउंट पन्नास हजार आहे",
"start":0.42,"end":3.18,"confidence":null}
{"type":"transcript.final","turn":0,"text":"Ho sir, flat cha booking amount 50 hazar aahe.",
"start":0.42,"end":3.18,"script":"latin","lane":"indic","repair_ms":1180,"fallback":null,"script_honoured":true}
{"type":"session.ended","session_id":"6f1c…","audio_seconds":124.37,"turns":31,"billed_seconds":125,
"cost":69.4445,"cost_currency":"INR_paise","cost_known":true}
| Event |
Fields |
session.started |
session_id, model, lane, language, script, encoding, sample_rate, endpointing_ms, vad |
speech.started |
turn (from 0), start — seconds into the audio you have sent, not wall-clock time |
transcript.partial |
turn, text, start, stage. stage: "streaming" — intl only, repeated while the turn is spoken, each one the whole turn so far. stage: "draft" — every lane, exactly once, when the turn ends; it adds end and confidence (null when not measured) |
transcript.final |
turn, text, start, end, script, lane, repair_ms, fallback, script_honoured |
session.ended |
session_id, audio_seconds, turns, billed_seconds, cost, cost_currency: "INR_paise", cost_known |
error |
code, message, fatal — after a fatal error the session ends and the socket closes |
pong |
— |
⚠️ Every speech.started gets exactly one transcript.final, and finals arrive in turn order. A later turn
can be ready first; it is held until the turn before it is sent, so you never have to reorder. A turn that held
no words — a cough, a line click — still gets its final, with text: "" and fallback: null.
⚠️ The draft is fast and rough; the final is the transcript. Use the draft to react while the caller is
still on the line, and store the final. fallback: "draft" means the check could not complete in time and the
final is the draft — you still get a final for the turn, never silence. fallback: "unavailable" means the
session ended on our side before the turn could be transcribed at all: text is "", but someone may have spoken.
⚠️ script applies to the final, not the draft. An Indian-language draft is written as it was first heard —
usually in the language's own alphabet, as in the example above — even when you asked for latin. Only the final
is converted, and only the final is checked.
⚠️ script_honoured is checked, not assumed. true means the final is in the alphabet you asked for, false
that it is not; null means the final has no letters to check — an empty turn.
How a session ends
| Close code |
Meaning |
| 1000 |
normal, after session.ended |
| 1009 |
a frame over 64 KiB |
| 1011 |
the gateway failed internally |
| 1012 |
the server is restarting. session.ended is sent first; reconnect, passing the last finals as context |
| 1013 |
transcription (model_unavailable) or usage recording (billing_unavailable) became unavailable mid-session. Reconnect shortly |
| 4401 |
key_revoked — the key was revoked during the session |
| 4402 |
insufficient_credit — the balance ran out, or key_limit_reached — this key's limit did, during the session. The seconds already received are charged |
| 4403 |
account_suspended — the account was suspended during the session |
| 4408 |
idle_timeout — no message from you for 60 seconds. Send audio, or ping while a vad=client call is on hold |
| 4413 |
session_too_long — 3 hours, or 3 hours of audio. Open a new session to continue |
The server sends a WebSocket ping every 20 seconds, which your client library answers for you, so a quiet line
is not dropped by a proxy on the way. It does not reset the 60-second idle rule, which counts messages from you.
Billed per second of audio you send
₹20/hour — lark-mini's own price — charged on every second of audio the session received, rounded up
once per session, not per turn. A 124.37-second call is 125 seconds, 69.4445 paise.
⚠️ Streaming and upload are separate lanes, and their prices may come to differ. Today the live
price is the upload price. Hindi live always uses the Indic lane's final pass, because live audio is 8 or
16 kHz by construction — the same rule the upload endpoint applies to telephone-rate Hindi. When the live
price changes it changes here and in the changelog first.
⚠️ Silence you send is audio you sent. Ten seconds of a muted line is ten billed seconds with no turn in it.
Stop streaming while a call is on hold, and send ping to keep the session.
⚠️ A long session is debited as it runs, not at the end. Every 60 seconds of audio the seconds so far are
charged, and your balance and the key's limit are checked again. When the balance is exhausted you get
error insufficient_credit, and when the key's limit is, error key_limit_reached — both close 4402, with the
seconds received so far charged. A session cannot run past the credit that pays for it.
What this endpoint refuses
Before the upgrade, as the usual JSON error envelope:
| Status |
Code |
When |
| 400 |
live_not_available |
model is anything other than lark-mini |
| 400 |
language_required |
no language |
| 400 |
language_not_supported |
a language not in the list above |
| 400 |
script_required |
an Indian language with no script |
| 400 |
unknown_script |
a script that is not latin or native or one of their synonyms — auto included |
| 400 |
unsupported_encoding |
an encoding other than linear16, mulaw, alaw |
| 400 |
unsupported_sample_rate |
not 8000 or 16000, or 16000 with mulaw/alaw |
| 400 |
invalid_endpointing |
endpointing_ms outside 300–1000 |
| 400 |
invalid_vad |
vad other than server or client |
| 400 |
domain_too_long / vocabulary_too_long |
over the limits above |
| 400 |
invalid_verify_vocabulary |
verify_vocabulary is not true or false |
| 400 |
context_too_long |
more than 3 context values, or one over 500 characters |
| 400 |
diarize_not_available_live |
diarize or speakers was sent |
| 401 |
invalid_api_key / key_revoked |
as on every endpoint |
| 402 |
insufficient_credit / key_limit_reached |
as on every endpoint |
| 403 |
account_suspended |
as on every endpoint |
| 426 |
upgrade_required |
a plain HTTP request, not a WebSocket upgrade |
| 429 |
rate_limited |
live capacity is full right now, or this account already has as many live sessions open as it may. Retry shortly |
| 503 |
tier_not_deployed |
live is not enabled for this account, or not serving on this deployment |
GET /v1/agent/live — Osprey Live
Live — Hindi and English, on every account (since 0.46.0, 2026-09-17). osprey-live is listed in
GET /v1/models with its caller languages.
A voice agent on one WebSocket. You stream the caller's audio in. MiniCrow hears each turn (Lark Live), decides
what to say and which of your tools to call (the Osprey Live brain, following your instructions), and speaks
the reply (Pica) back down the same socket while it is still being written. You describe the agent once, in
session.configure. You run your own tools. MiniCrow runs the three parts and bills each one.
About ₹6.50 per 1 million tokens — means around ₹66 per hour including STT + LLM + TTS (best for AI call agents). An illustration for a typical voice agent with prompt caching, not a quote: the brain is billed per token type, speech-to-text per second of audio and text-to-speech per character.
Languages. The caller's language is the language query parameter.
| Language |
Code |
Status |
| Hindi |
hi |
Available |
| English |
en |
Available |
| Marathi |
mr |
Coming soon |
| Tamil |
ta |
Coming soon |
| Telugu |
te |
Coming soon |
| Every other language |
|
Coming soon |
A language that is coming soon is refused with 400 unsupported_language and a message saying it is coming soon.
wss://api.minicrow.com/v1/agent/live?model=osprey-live&language=hi&encoding=mulaw&sample_rate=8000&output=mulaw_8k&voice=david
Authorization: Bearer mc_...
⚠️ The key goes in the Authorization header of the upgrade, and nowhere else. Not in the URL, and not in a
message. A browser's WebSocket cannot set a header, so connect from your server: Osprey Live is for server-side
clients, such as the relay between your telephony provider and MiniCrow.
Python
# pip install "websockets>=14"
import asyncio, json, os
from urllib.parse import urlencode
import websockets
params = urlencode({
"model": "osprey-live", "language": "hi", "encoding": "mulaw", "sample_rate": 8000,
"output": "mulaw_8k", "voice": "david",
})
URL = f"wss://api.minicrow.com/v1/agent/live?{params}"
HEADERS = {"Authorization": f"Bearer {os.environ['MINICROW_API_KEY']}"}
with open("agent.json", encoding="utf-8") as f: # your agent spec, see "The agent spec" below
AGENT = json.load(f)
async def run_tool(ws, call):
"""Your own code: look something up, check a slot, send a message. The socket keeps running meanwhile."""
if call["name"] == "check_appointment_slots":
output, failed = {"doctor": call["arguments"]["doctor"], "available": ["10:40", "11:20"]}, False
else:
output, failed = {"error": "unknown tool"}, True
await ws.send(json.dumps({"type": "tool.result", "call_id": call["call_id"], "output": output,
"is_error": failed}))
async def send_audio(ws, path):
with open(path, "rb") as audio: # raw 8 kHz mu-law, the caller only, no header
while chunk := audio.read(160): # 160 bytes = 20 ms
await ws.send(chunk) # bytes are sent as a binary frame
await asyncio.sleep(0.02) # real-time pace, as a live call arrives
await ws.send(json.dumps({"type": "end"}))
async def main():
tasks = []
try:
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
with open("agent-reply.ulaw", "wb") as speaker:
async for message in ws: # ends when the server closes
if isinstance(message, bytes): # the agent's voice, in the `output` format
speaker.write(message)
continue
event = json.loads(message)
kind = event["type"]
if kind == "session.started":
await ws.send(json.dumps({"type": "session.configure", "agent": AGENT}))
elif kind == "session.configured":
for warning in event["warnings"]:
print("spec warning:", warning)
tasks.append(asyncio.create_task(send_audio(ws, "caller.ulaw")))
elif kind == "response.audio.started":
print("agent:", event["text"])
elif kind == "tool.call":
tasks.append(asyncio.create_task(run_tool(ws, event)))
elif kind == "turn.usage":
print(f"turn {event['turn']}: {event['cost']} paise")
elif kind == "session.ended":
print(f"session: {event['cost']} paise, reason {event['reason']}")
elif kind == "error":
print("error:", event["code"], event["message"])
except websockets.exceptions.InvalidStatus as refused: # refused before the socket opened
print(refused.response.status_code, refused.response.body.decode())
finally:
for task in tasks:
task.cancel()
asyncio.run(main())
JavaScript (Node)
// npm install ws — save as agent.mjs, run with node agent.mjs
import fs from "node:fs";
import WebSocket from "ws";
const params = new URLSearchParams({
model: "osprey-live", language: "hi", encoding: "linear16", sample_rate: "16000", output: "pcm16_24k",
});
const ws = new WebSocket(`wss://api.minicrow.com/v1/agent/live?${params}`, {
headers: { Authorization: `Bearer ${process.env.MINICROW_API_KEY}` },
});
const agent = JSON.parse(fs.readFileSync("agent.json", "utf8"));
const speaker = fs.createWriteStream("agent-reply.pcm"); // raw 24 kHz 16-bit PCM, no header
// Refused before the socket opened: the body is the usual JSON error.
ws.on("unexpected-response", (req, res) => {
let body = "";
res.on("data", (chunk) => (body += chunk));
res.on("end", () => console.error(res.statusCode, body));
});
async function sendAudio() {
const audio = fs.readFileSync("caller-16k.pcm"); // raw 16 kHz 16-bit PCM, the caller only
for (let i = 0; i < audio.length && ws.readyState === WebSocket.OPEN; i += 3200) { // 3200 bytes = 100 ms
ws.send(audio.subarray(i, i + 3200));
await new Promise((resolve) => setTimeout(resolve, 100));
}
if (ws.readyState === WebSocket.OPEN) ws.send(JSON.stringify({ type: "end" }));
}
ws.on("message", (data, isBinary) => {
if (isBinary) return speaker.write(data); // the agent's voice
const event = JSON.parse(data.toString());
switch (event.type) {
case "session.started":
ws.send(JSON.stringify({ type: "session.configure", agent, tool_timeout_ms: 8000 }));
break;
case "session.configured":
sendAudio();
break;
case "tool.call": {
const output = { status: "sent" }; // run your tool here
ws.send(JSON.stringify({ type: "tool.result", call_id: event.call_id, output, is_error: false }));
break;
}
case "response.audio.started":
console.log("agent:", event.text);
break;
case "turn.usage":
case "session.ended":
console.log(event.type, event.cost, "paise");
break;
case "error":
console.error("error:", event.code, event.message);
break;
}
});
ws.on("close", (code, reason) => {
speaker.end();
console.log("closed", code, reason.toString());
});
Query
Everything about the audio is decided before the socket opens. A wrong parameter is an ordinary JSON error on the
upgrade request, not a socket that opens and then closes.
| Parameter |
Values |
|
model |
osprey-live |
required |
language |
required. hi · en |
the language the caller speaks. Other languages are coming soon: they are refused with unsupported_language, never guessed |
encoding |
required. linear16 · mulaw · alaw |
the caller's audio. linear16 is signed 16-bit little-endian mono PCM; μ-law and A-law are G.711, as a telephone line carries them |
sample_rate |
required. 8000 · 16000 |
mulaw and alaw are 8000 only |
output |
mulaw_8k · pcm16_24k |
the agent's voice. Omitted, it is mulaw_8k when encoding=mulaw and pcm16_24k otherwise. pcm16_24k is raw signed 16-bit little-endian mono PCM at 24 kHz, with no header |
voice |
a built-in voice or one of yours (mcv_…) |
GET /v1/audio/voices lists both. Omitted, it is david when the agent speaks Hindi and robert when it speaks English |
endpointing_ms |
300–1000, default 400 |
how much silence ends the caller's turn. Lower answers sooner and splits more sentences in two |
vad |
server (default) · client |
client: you decide where the caller's turn ends by sending {"type":"turn_end"} |
barge_in |
draft (default) · speech · off |
what happens when the caller talks over the agent. See Interruptions |
transcripts |
false (default) · true |
true sends you each caller turn as input.transcript |
Opening a session
- The upgrade succeeds, and you receive
session.started.
- Within 10 seconds, and before any audio, you send
session.configure with your agent spec.
- You receive
session.configured: the spec was accepted. It lists your tools, the limits in force and any
warnings about the spec. Start sending audio now.
Audio sent before session.configured is dropped with a non-fatal error not_configured. No session.configure
within 10 seconds closes the session with config_timeout (4400).
The agent spec
session.configure carries everything the agent knows: who it is, how it speaks, your instructions, your tools and
a few short invented examples. MiniCrow adds its own rules on top of it (how to read a fast draft of the caller's turn,
how to behave mid-call, when to use tools) that are the same for every customer. The Osprey Live agent guide covers
every field and how to write the parts that decide whether a tool gets called.
{"type":"session.configure",
"agent":{
"spec_version":"osprey-live-prompt/0.1",
"persona":{"name":"Asha"},
"caller":{"pronoun":"they"},
"languages":{"caller_speaks":["hi","en"],"agent_speaks":"hi","agent_script":"latin"},
"speech":{"reprompt":"Ji, boliye?","greeting_words":["hello","namaste"],
"never_claim_before_result":["booked","sent"],
"block_before_result":["book ho gaya","sms bhej diya"]},
"brief":"You are Asha, the appointment desk voice for Leafview Family Clinic, Indore. You are on a live phone call and have already greeted the caller. Speak one short sentence of romanised Hindi per reply. To check or book an appointment, call the tool.",
"tools":[{"name":"check_appointment_slots","kind":"booking",
"description":"Check which appointment times are free for a doctor on a date. Commits nothing.",
"when_to_call":"CALL THIS BEFORE saying any appointment date or time.",
"when_not_to_call":"Do not call it again for the same doctor and date in the same call.",
"acknowledgement":"Ji, time dekh leti hoon.",
"reminder_trigger":"an appointment time",
"parameters":{"type":"object","properties":{"doctor":{"type":"string"},"date":{"type":"string"}},
"required":["doctor","date"]}},
{"name":"send_sms_confirmation","kind":"action_needs_confirmation",
"description":"Send an SMS with the appointment details.",
"when_to_call":"CALL THIS right after the caller says yes to an SMS.",
"when_not_to_call":"Do not call it again for the same message on the same call.",
"acknowledgement":"Ji, SMS bhej rahi hoon.",
"parameters":{"type":"object","properties":{"mobile":{"type":"string"}},"required":["mobile"]},
"confirm":"runtime"}],
"examples":[{"history":[{"speaker":"agent","text":"Ji, consultation fee 500 rupees hai."},
{"speaker":"customer","text":"Theek hai"}],
"customer_draft":"कल सुबह दिखाना है",
"acknowledgement":"Ji, time dekh leti hoon.",
"tool_call":{"name":"check_appointment_slots","arguments":{"doctor":"Dr. Arjun Rao","date":"2026-03-04"}},
"tool_result":{"available":["10:40"]},
"reply":"Kal subah 10:40 par Dr. Rao ka time khali hai ji, 10:40 theek rahega?"}]},
"tool_timeout_ms":8000,
"history":[{"speaker":"agent","text":"Namaste, main Asha bol rahi hoon Leafview Family Clinic se."}]}
| Field |
|
agent |
required. The agent spec. spec_version is "osprey-live-prompt/0.1" |
tool_timeout_ms |
1000–30000, default 8000. How long the agent waits for a tool.result |
history |
at most 64 lines, oldest first, each {"speaker":"agent","text":"…"} or {"speaker":"customer","text":"…"}: what was already said on this call. Use it for a greeting you already played, and when you reconnect |
Only in the spec, and never in what the agent is told:
| Field |
|
languages.agent_speaks |
hi or en: the language of the agent's voice. The voice you chose must speak it |
speech.block_before_result |
extra phrases the agent may never say before a tool has returned a result, on top of MiniCrow's own list (see What the runtime enforces). List only past-tense outcome phrases — "book ho gaya", "has been sent" — never ordinary words such as "available" |
tools[].confirm |
model (default) or runtime. See Tools |
hearing.vocabulary |
0.41.0. Names your callers say that a general listener would not spell right — people, products, places, codes — as the transcript should show them, up to 100. Each entry is a string or {"term": "Dr. Arjun Rao", "spoken": ["doctor arjun rao", "डॉक्टर अर्जुन राव"]} with up to 3 spoken forms. See Hearing vocabulary |
tools[].execution |
{"type":"client"}, the default: you run the tool when you receive tool.call |
⚠️ Everything in examples must be invented. The agent copies from examples: names, numbers and dates in them
reach real replies. session.configured.warnings flags digit runs that look real and numbers shared with your brief.
Hearing vocabulary
The caller's turn reaches the agent as a fast draft, and a fast draft hears a name it has never seen as the nearest
everyday words: "Orbit" as "और भी", "Zenith" as "जैने", "WhatsApp" as "what's that mean". Tell the hearing what to listen
for and it is boosted while it decodes — the standard hotwords of speech APIs — at no extra cost and about 20 ms.
"hearing": {
"vocabulary": [
"Orbit", "Zenith",
{"term": "Dr. Arjun Rao", "spoken": ["doctor arjun rao", "डॉक्टर अर्जुन राव"]},
{"term": "Tarangpur", "spoken": ["तरंगपुर"]}
]
}
- List names, not words. People, products, places, models, codes. An everyday word ("today", "road", "आज") is
never boosted, however you write it; boosting common words made the hearing hear them where they were not said.
- Spell the term as the transcript should show it. That spelling is what the agent reads and what a tool receives.
- For Hindi and Marathi callers, add the Devanagari form in
spoken. The hearing writes those languages in
Devanagari, so "Orbit" alone cannot match; {"term": "Orbit", "spoken": ["ऑर्बिट"]} can. English callers need no
spoken form. Add a spoken form too when callers say a name differently from how it is written.
- Your tools'
enum values are boosted on their own (skin_specialist is heard as "skin specialist"); do not list them.
- Limits: 100 entries; a term at most 5 words and 64 characters; 3 spoken forms of at most 64 characters. Over any of
these the spec is refused (
invalid_agent) with the entry named.
- Names the vocabulary does not carry — the caller's own town, a person the spec never mentions — are not helped;
for those the runtime read-back (
confirm: "runtime") is the guard.
- A yes to the read-back sends the held call (0.43.2): after
confirm: "runtime" reads the details back, the
caller's "haan", "ho", "बरोबर", "yes" makes the agent's next reply call the tool with the same details; a changed
value is read back again. Before 0.43.2 the agent could ask "shall I send it?" a second time instead.
- A phone number the caller never said is never sent (0.43.1): a placeholder such as 9876543210, or a number that
appears only in your examples, in any tool's phone argument is held; the agent is told to use
caller_id when the caller
said to use the number they are calling from, and to ask for the number otherwise. Measured: the default brain had
written 9876543210 — which appears nowhere in the spec — into 6 of 99 tool turns.
- On for Hindi, Marathi and English — the languages it is measured on. Hindi and Marathi boost at one strength,
English at a slightly lower one: on 88 English turns the boosted hearing made fewer word errors than the plain decode
(12.4 % against 14.3 %), caught every covered name (13 of 13 against 12) and added one word that was not said. The
other 18 languages use the vocabulary as a spelling hint only, without the boost, until each is measured the same way.
- Hindi and Marathi hearing is locked to Devanagari (0.43.0), with or without a vocabulary. The fast draft is
written without knowing the language, and 22 of 166 Hindi and Marathi turns had come back with letters of another
Indian script — a name in Bengali letters matches no Devanagari spelling. With the lock: 0 of 166; English words in
English letters, digits and punctuation are untouched.
Measured (replay, 166 Hindi and Marathi turns, invented clinic / car service / courier calls plus real Marathi calls): of
the spoken names the vocabulary covered, the boosted hearing caught 5 of 7 against 2 of 7 in production, with the word
error of the whole set unchanged; with the Devanagari lock as well, 6 of 7 and the word error 0.7 points lower (within
the run's noise). The required tool calls of the same brain did not move beyond the replay's run-to-run noise.
Writing a spec that calls its tools
MiniCrow's own rules name no business. Whether the agent calls your tool at the right moment depends on your spec, and
session.configured.warnings checks it against these rules. A warning never refuses a spec.
| Warning |
The rule |
no when_to_call |
Every tool says when to call it, starting with "CALL THIS" and naming situations, not your test sentences. A tool described only by what it does was not called |
N required fields |
A lookup or booking with more than 3 required fields waits until the caller has answered all of them. Require only what the tool cannot run without; ask the rest after the result |
when_to_call waits for a confirmation |
A lookup or booking commits nothing, so it should not wait for a yes or for a "confirmed" value. Write "call as soon as <required fields> are known" |
is dictated by the caller; set "confirm":"runtime" |
An action whose arguments include a phone number, name, email, account or order number the caller says aloud. With runtime confirmation the values are read back, numbers digit by digit, and a misheard name or number is asked again or spelt before the call reaches you |
should say when NOT to call it |
An action says when not to call it: already done on this call, a number still being dictated |
no example calls it |
A lookup or booking that no example calls. One invented example that calls it at the right moment teaches more than a rule |
no examples |
A spec without examples made the fewest required calls |
names X, which is not one of your tools |
The brief names a tool that is not in tools, and the agent tries to call it |
acknowledgement is over 6 words |
The spoken line before a tool is 3 to 5 words and promises nothing |
looks like a phone or account number |
See above: examples are invented |
Describe each argument as a value and where it comes from ("the phone number to send to: the caller's own number if
they say to use it, otherwise the number they give"), not as steps the caller goes through ("the number they dictated
and confirmed"): the second wording made the agent ask callers to dictate a number they had just told it to use.
A spec that cannot be used ends the session with error invalid_agent, fatal: true, and close 4400. The
message names the field and the reason, for example tools[1].name: say_filler is reserved. It happens when:
- a required field is missing, a tool name is repeated, a
kind is not lookup, booking,
action_needs_confirmation or other, or an example calls a tool with arguments its parameters do not allow;
- the brief is over 24,000 characters, or there are more than 8 tools or more than 12 examples;
agent_speaks is not hi or en, or the voice does not speak it;
- the spec contains a
cache_control key anywhere, or an example carries a role;
- a tool is named
say_filler (reserved), or execution is anything but {"type":"client"}.
Keys the format does not use are ignored: any key starting with _, and examples[].id. session.configure may be
up to 256 KiB. A second session.configure is a non-fatal already_configured.
What you send
| Message |
|
| binary frame |
the caller's audio in the encoding and sample_rate you declared, 20 ms to 1 s per frame. A frame over 64 KiB closes the session with 1009 |
session.configure |
once, first. See above |
tool.result |
the answer to a tool.call: {"type":"tool.result","call_id":"call_7f3aQ2mX9kLp","output":{"status":"sent"},"is_error":false}. output is any JSON value or a string, at most 32 KiB |
response.say |
speak this text as it is, with no brain call: {"type":"response.say","text":"Namaste, main Asha bol rahi hoon Leafview Family Clinic se. Kya do minute baat kar sakte hain?"}. At most 500 characters. While a response is running it is refused with a non-fatal busy |
response.cancel |
stop the current response, for example on a keypad press: {"type":"response.cancel"} |
response.played |
optional. How much of a cancelled response the caller really heard: {"type":"response.played","response_id":"r_12","played_ms":1840} |
turn_end |
vad=client only: the caller's turn ends here |
end |
finish. A running response completes (at most 15 seconds), the last usage is charged, then session.ended and close 1000 |
ping |
answered with {"type":"pong"} |
input.image |
0.68.0. a picture from the caller's camera: {"type":"input.image","data":"<base64 JPEG, PNG or WebP>"}, at most 192 KB. Send one every 1–3 seconds while the camera is on. See Vision |
Any other text message is answered with a non-fatal error bad_message, and the session carries on.
⚠️ Send only the caller's audio, on one channel. If the agent's own voice comes back in your input (a mixed
recording, a speakerphone, a call leg that carries both parties), the agent hears itself talking and stops
mid-sentence. Take the caller's inbound track only.
What you receive
Text frames are JSON events. Binary frames are the agent's voice in your output format, at most 200 ms each. One
turn with a tool, in order, up to the caller's next turn (which closes turn 7's charge):
{"type":"session.started","session_id":"6f1c2a9e-3b7d-4c1a-9f0e-5d8b7a6c4e21","model":"osprey-live"}
{"type":"session.configured","session_id":"6f1c2a9e-3b7d-4c1a-9f0e-5d8b7a6c4e21","model":"osprey-live","input":{"encoding":"mulaw","sample_rate":8000},"output":{"format":"mulaw_8k","sample_rate":8000},"voice":"david","language":"hi","tools":["check_appointment_slots","send_sms_confirmation"],"limits":{"max_seconds":10800,"tool_timeout_ms":8000,"max_passes":3},"warnings":[]}
{"type":"input.speech_started","turn":7,"start":41.232}
{"type":"input.transcript","turn":7,"stage":"draft","text":"कल सुबह डॉक्टर राव का टाइम है क्या","start":41.232,"end":43.02}
{"type":"response.started","response_id":"r_12","turn":7,"kind":"reply"}
{"type":"response.audio.started","response_id":"r_12","clause":0,"kind":"speech","text":"Ji, time dekh leti hoon.","format":"mulaw_8k"}
{"type":"tool.call","response_id":"r_12","turn":7,"call_id":"call_7f3aQ2mX9kLp","name":"check_appointment_slots","arguments":{"doctor":"Dr. Arjun Rao","date":"2026-03-04"}}
{"type":"response.audio.done","response_id":"r_12","clause":0,"audio_ms":1520}
{"type":"response.audio.started","response_id":"r_12","clause":1,"kind":"speech","text":"Kal subah 10:40 par Dr. Rao ka time khali hai ji, 10:40 theek rahega?","format":"mulaw_8k"}
{"type":"response.audio.done","response_id":"r_12","clause":1,"audio_ms":3880}
{"type":"response.done","response_id":"r_12","turn":7,"status":"completed","passes":2,"tool_calls":1,"sent_ms":5400}
{"type":"input.speech_started","turn":8,"start":52.108}
{"type":"input.transcript","turn":8,"stage":"draft","text":"हां 10:40 ठीक है","start":52.108,"end":53.3}
{"type":"turn.usage","turn":7,"brain":{"input_tokens":412,"cached_input_tokens":4220,"cache_write_tokens":0,"output_tokens":61,"audio_input_tokens":0,"passes":2,"cost":3.3},"voice":{"characters":93,"cost":8.37},"cost":11.67,"cost_currency":"INR_paise","cost_known":true}
Binary frames arrive between each response.audio.started and its response.audio.done. When the agent speaks an
acknowledgement before a tool, tool.call arrives right after the acknowledgement's first binary frame, so you can
start the tool while the caller hears it.
| Event |
Fields |
session.started |
session_id, model |
session.configured |
session_id, model, input (encoding, sample_rate), output (format, sample_rate), voice, language, tools (names), limits (max_seconds, tool_timeout_ms, max_passes), vision on a session that takes pictures (formats, max_image_bytes, min_interval_ms, frames_per_turn), warnings (always an array) |
input.speech_started |
turn (from 0), start: seconds into the audio you have sent |
input.transcript |
transcripts=true only. turn, stage: "draft", text, start, end. A fast draft, written as it was first heard: usually in the language's own script |
response.started |
response_id, turn, kind: reply (the agent answering a turn) or say (your response.say) |
response.audio.started |
response_id, clause (from 0), kind: speech or filler, text: exactly the words in the audio that follows, format |
response.audio.done |
response_id, clause, audio_ms |
tool.call |
response_id, turn, call_id, name, arguments (an object that matches the tool's parameters) |
tool.timeout |
call_id, after_ms: no result arrived in time, and the agent carried on without it |
response.done |
response_id, turn, status, passes, tool_calls, sent_ms, and reason when status is cancelled. status: completed · cancelled (reason: barge_in or client) · incomplete (the reply stopped part-way) · failed |
turn.usage |
the charge for one turn. See Billing |
error |
code, message, fatal. After a fatal error the session ends: session.ended, then the close |
session.ended |
session_id, reason, turns, totals per part and cost. See Billing |
pong |
— |
⚠️ A binary frame belongs to the clause of the latest response.audio.started that has no response.audio.done
yet. One writer sends every frame, text and audio, so the order on the socket is the order to play.
⚠️ The audio arrives at the pace it should be played, plus a lead of at most 200 ms. Play it as it arrives; you
do not need a large buffer, and when the agent stops (an interruption), it goes quiet within about 200 ms without
you clearing anything.
A turn, step by step
- The caller speaks; you receive
input.speech_started.
- The caller stops. After
endpointing_ms of silence the turn is drafted.
- The brain reads your spec, the call so far and the draft, and starts writing. The first clause goes to the voice
at once, and later clauses follow while the brain is still writing. You get
response.started, then each
clause's response.audio.started, its audio, and response.audio.done.
- If the brain calls a tool, it usually speaks your tool's
acknowledgement first ("Ji, time dekh leti hoon."),
then you receive tool.call. When your tool.result arrives, the brain is asked again with the result, and it
speaks the answer. A reply may take up to 3 passes.
response.done. What the caller heard goes into the call's history.
turn.usage arrives once the turn's charge is recorded: after the caller's next turn is drafted, or when the
session ends.
The first turns of a call are slower by a few seconds, while your instructions and tools are stored in the prompt
cache; later turns read them from it. If no audio has gone out 1.8 seconds after the caller's turn was drafted, or
a tool is slow, the agent says a short holding phrase in the same voice ("Ji, ek second."). It arrives as a clause with
kind: "filler", it is billed as voice characters, and its words enter the history like any other speech.
A sound under 0.3 seconds (a cough, a "hm") is not a turn and gets no reply. A caller turn is merged with the next
one, and not answered alone, when it was cut at 20 seconds of continuous speech, or when the caller started speaking
again before the agent's first audio went out. A turn with no clear words gets your speech.reprompt; after two
reprompts in a row, a third empty turn gets no reply.
Vision: the caller's camera
0.68.0. Osprey Live sees what the caller's camera sees — smart glasses, a phone, a robot, a kiosk, a car. Send a
picture every one to three seconds as input.image, on the same socket as the audio. When the caller finishes a
turn, the reply is written with the pictures taken while they spoke — from two seconds before they started, at most
three, spread over the turn with the newest always among them — beside their words. So "what is this?" is answered
about what the camera showed when it was asked.
{"type":"input.image","data":"/9j/4AAQSkZJRgABAQEASABIAAD…"}
async def camera(ws, frames): # frames: JPEG bytes from your camera, as they come
async for jpeg in frames:
await ws.send(json.dumps({"type": "input.image", "data": base64.b64encode(jpeg).decode()}))
await asyncio.sleep(1.5)
|
|
| Formats |
JPEG, PNG or WebP, base64. A data:image/jpeg;base64,… URL is taken too |
| Size |
at most 192 KB a picture; 640 to 1024 pixels wide is plenty |
| How often |
one every 1–3 seconds while the camera is on. At most one a second is kept; a faster camera's extra pictures are dropped |
| Per turn |
at most 3 pictures go with a reply. A turn shorter than your interval takes the newest picture while it is at most 5 seconds old |
| What is seen |
only the turn's own pictures. What the agent said about a picture stays in the call's history as text; earlier pictures are not sent again |
A session that takes pictures says so in session.configured:
"vision":{"formats":["jpeg","png","webp"],"max_image_bytes":196608,"min_interval_ms":1000,"frames_per_turn":3}.
Send pictures after it arrives. A picture that is too large or not JPEG, PNG or WebP — or any picture sent to a
session without vision — is answered with a non-fatal error bad_message, and the call goes on.
Billing. Pictures are brain input tokens of the turn they go with: about 1,100 tokens a picture, so a turn with
three pictures has about 3,300 more input_tokens in its turn.usage — about 9 paise. A turn without pictures costs
what it did.
- You run every tool. MiniCrow never calls your systems: it sends
tool.call, and you answer with tool.result
carrying the same call_id.
arguments are checked against the tool's parameters before you see them. A call that does not match is not
sent to you; the agent is told what was wrong and carries on.
- Calls of one pass are sent together, at most 2 per pass, and the agent waits for all of their results.
- No result within
tool_timeout_ms: you receive tool.timeout, and the agent is told the tool timed out. A result
that arrives later is still recorded in the history, because the action may have happened, but that reply does not
change.
output goes to the agent trimmed to 4 KiB, and into the history trimmed to 300 characters. Send short JSON that
holds only what the agent may say.
is_error: true, a top-level error key, or status failed or error in output count as a failed result.
- A
call_id that is unknown, already answered or timed out gets a non-fatal unknown_call_id. An output over
32 KiB gets result_too_large.
Two confirmation modes, per tool:
confirm |
|
model (default) |
the brain decides when the caller has agreed, following your when_to_call and when_not_to_call |
runtime |
the first call is not sent to you. The agent is told to read the arguments back to the caller, numbers digit by digit, and to ask the caller to spell or repeat a name, email or number that may have been misheard, then call again with the corrected value. If it calls the same tool with the same arguments within the next 2 caller turns and the caller's turn is a yes ("haan", "ho", "yes", "sahi hai"; not a question such as "bheja?", not a "no"), you receive tool.call. Otherwise the call keeps waiting and the agent is told to answer the caller and read back again. Different arguments replace the waiting call and need their own yes. A phone number that does not have 10 digits (after +91 or a leading 0) is never sent; the agent is told to ask for the number again. Identifiers such as contact_id and 1800/1860 toll-free numbers are not checked. It costs at least one extra pass |
Interruptions
The caller talks over the agent when input.speech_started arrives while a response is playing.
barge_in |
|
draft (default) |
the agent's audio holds at once. If the caller's sound turns out to be under 0.3 seconds (a cough), or a one- or two-word acknowledgement that is not a question ("haan", "ho", "achha", "hmm", "ok", "theek hai"), the held audio resumes where it stopped. Anything else is real speech: the response is cancelled and the caller's words become the next turn. A yes said over a runtime read-back, or over a question the caller has already heard, is kept as the caller's next turn |
speech |
the response is cancelled as soon as the caller starts speaking |
off |
the agent finishes; the caller's turn is answered afterwards |
A cancelled response ends with response.done status: "cancelled", reason: "barge_in", and sent_ms.
response.cancel does the same with reason: "client". Tool calls already sent stay valid, and their results are
still recorded; calls not yet sent are dropped.
Only what the caller heard enters the history. By default, that is every clause that finished playing more than
200 ms before the cancel. If you send response.played within 1 second of the cancel, your played_ms decides
instead. A clause the caller heard only in part is left out; it is never cut mid-word.
Opening lines and reconnects
- Outbound calls (the agent speaks first): send
response.say with the greeting after session.configured, or
play your own greeting and pass it in history.
- Inbound calls (the caller speaks first): just start sending audio. The agent answers the first turn.
- Reconnects: on close 1012, 1013 or a dropped connection, open a new session and send the call so far in
history. The agent continues the call and does not greet again.
response.say text is spoken exactly as written, enters the history as the agent's line, and is billed as voice
characters.
What the runtime enforces
These hold whatever your instructions, the caller or a tool result say:
- No claimed outcome without a result. A clause such as "book ho gaya" or "has been sent" is not spoken unless
a tool returned a successful result earlier in the same response, or a
booking or action_needs_confirmation
tool succeeded earlier in the call. MiniCrow's list covers common Hindi (Latin and Devanagari) and English outcome
phrases; speech.block_before_result adds yours. A promise of what the agent is about to do ("bhej rahi hoon") is
allowed. If nothing is left to say, the agent says your speech.reprompt.
- No second greeting. After the agent's first line (audio the caller heard counts, even when it was
interrupted), a clause that starts with a greeting, in any of the supported scripts, has the greeting removed, and
a clause that only re-introduces the agent ("Main Asha bol rahi hoon, …") is dropped. Your
speech.reprompt is
never changed.
- No repeated identical action. An
action_needs_confirmation call with the same arguments as one that already
succeeded on this call is not sent to you again; the agent is told it is already done.
- No spoken tool talk. A reply never reads out a tool call the agent wrote into its words instead of making
it — JSON, a function call or a snake_case name. Those words are taken out before the voice; the rest is spoken.
- Limits per caller turn: at most 3 passes, at most 2 tool calls per pass, and at most 600 characters spoken
per response.
A withheld clause is not spoken, not billed and not put in the history.
Languages and voices
|
|
| The caller speaks |
hi (Hindi) or en (English): the language of the session. Other languages are coming soon |
| The agent speaks |
hi or en: agent.languages.agent_speaks |
| Voices |
any built-in voice or voice of yours that speaks agent_speaks. Defaults david (Hindi) and robert (English) |
Say in your brief how the agent writes: "romanised Hindi in Latin letters, never Devanagari", and how it says
numbers, prices and times. The voice reads what the agent writes.
Billing
Three parts, each billed on its own, in rupees, from your MiniCrow credits. Prices are fixed when a session opens.
About ₹6.50 per 1 million tokens — means around ₹66 per hour including STT + LLM + TTS (best for AI call agents). An illustration for a typical voice agent with prompt caching, not a quote: the brain is billed per token type, speech-to-text per second of audio and text-to-speech per character.
| Part |
Billed on |
Price |
| Hearing (Lark Live) |
every second of audio the session receives, silence included, rounded up once per session |
₹20 per hour |
| Brain (Osprey Live) |
the tokens of every pass, by type (below) |
per million tokens |
| Voice (Pica) |
characters spoken to the caller: replies, acknowledgements, holding phrases and response.say |
₹9 per 10,000 characters |
| Brain token type |
What it is |
₹ per million tokens |
| Input |
the part of the prompt not read from the cache: the latest lines, the caller's draft, tool results |
₹27.50 |
| Cached input |
the part read from the prompt cache: usually your brief, tools and examples |
₹2.75 |
| Cache write |
storing the fixed part in the cache, charged each time it happens, on top of cached input |
₹9.17 |
| Output |
what the brain writes: the spoken reply, tool calls and its reasoning |
₹165.00 |
audio_input_tokens in the usage events is always 0 on Osprey Live: the caller's audio is billed as hearing.
Every turn reports its charge once it is recorded:
{"type":"turn.usage","turn":4,
"brain":{"input_tokens":203,"cached_input_tokens":4220,"cache_write_tokens":0,"output_tokens":33,
"audio_input_tokens":0,"passes":1,"cost":2.2632},
"voice":{"characters":114,"cost":10.26},
"cost":12.5232,"cost_currency":"INR_paise","cost_known":true}
That turn read 4,220 tokens from the cache, so the brain cost 2.2632 paise; the 114 spoken characters cost 10.26
paise. The last event gives the totals:
{"type":"session.ended","session_id":"6f1c2a9e-3b7d-4c1a-9f0e-5d8b7a6c4e21","reason":"client_end","turns":12,
"hearing":{"audio_seconds":184.4,"billed_seconds":185,"cost":102.778},
"brain":{"input_tokens":2911,"cached_input_tokens":54860,"cache_write_tokens":4220,"output_tokens":602,
"audio_input_tokens":0,"cost":36.8931},
"voice":{"characters":1288,"cost":115.92},
"cost":255.5911,"cost_currency":"INR_paise","cost_known":true}
cost is always paise charged, to four decimals. The session's cost is exactly the sum of what was
debited: its turns and its hearing. Your dashboard shows the same rows.
cached_input_tokens includes tokens stored in the cache during that turn; cache_write_tokens is charged on
top of them.
passes is how many times the brain was asked in the turn. A tool call usually makes two.
- ⚠️ On Osprey Live,
cost_known: false means part of the usage could not be measured and was not charged, for
example when the connection dropped before a pass reported its tokens. cost is then a lower bound, never an
estimate. (On other endpoints cost_known keeps the meaning described there.)
- A cancelled response is still billed for the tokens the brain used and the characters that were sent to the
voice.
⚠️ Voice is the largest part of the bill — about half. An illustration from two test calls with a 4,200-token
spec, 336 caller turns an hour and about 107 spoken characters a turn: hearing ₹20.00 + brain ₹13.50 + voice ₹32.24
= ₹65.74 per call-hour (about ₹66). Short spoken replies lower it more than anything else. Keep the fixed part (brief, tools, examples)
identical for the whole call so it stays cached.
⚠️ The balance is checked while the session runs. Hearing is charged every 60 seconds of audio, and your balance
and the key's limit are checked before every turn and every minute. When credit runs out, a running response
finishes and is billed, then you get error insufficient_credit (or key_limit_reached), session.ended, and
close 4402.
How a session ends
| Close code |
Meaning |
| 1000 |
normal, after end and session.ended |
| 1009 |
an audio frame over 64 KiB, a session.configure over 256 KiB, or an input.image over 320 KiB |
| 1011 |
the gateway failed internally, or the agent could not answer 3 turns in a row (brain_unavailable) |
| 1012 |
the server is restarting (server_restarting). session.ended is sent first; reconnect with history |
| 1013 |
hearing, voice or the agent became unavailable mid-session (hearing_unavailable, voice_unavailable, service_unavailable), usage could not be recorded (billing_unavailable), or your client did not read fast enough (slow_consumer). Reconnect shortly with history |
| 4400 |
invalid_agent (the spec cannot be used) or config_timeout (no session.configure within 10 seconds) |
| 4401 |
key_revoked: the key was revoked during the session |
| 4402 |
insufficient_credit or key_limit_reached during the session. What was used so far is charged |
| 4403 |
account_suspended during the session |
| 4408 |
idle_timeout: no message or audio from you for 60 seconds |
| 4413 |
session_too_long: 3 hours. Open a new session with history to continue |
Non-fatal errors leave the session open: bad_message, not_configured, already_configured, busy,
unknown_call_id, result_too_large, say_too_long, voice_unavailable (part of a reply could not be spoken; it is
not billed and not put in the history), brain_unavailable (the agent could not answer this turn and said your
reprompt) and brain_interrupted (the reply stopped part-way; what was spoken stays).
The server sends a WebSocket ping every 20 seconds, which your client library answers for you. It does not reset the
60-second idle rule, which counts messages from you.
What this endpoint refuses
Before the upgrade, as the usual JSON error envelope:
| Status |
Code |
When |
| 400 |
model_not_found |
model is not osprey-live |
| 400 |
unsupported_language |
no language, or a language other than hi and en. For a language that is coming soon, the message says so |
| 400 |
invalid_parameter |
a missing or wrong encoding, sample_rate, output, voice, endpointing_ms, vad, barge_in or transcripts. The error names the param |
| 400 |
unsupported_parameter |
script, diarize, domain, vocabulary or context: parameters of the transcription socket that this one does not take |
| 401 |
invalid_api_key / key_revoked |
as on every endpoint |
| 402 |
insufficient_credit / key_limit_reached |
as on every endpoint |
| 403 |
account_suspended |
as on every endpoint |
| 404 |
not_found |
Osprey Live is not enabled for this account |
| 426 |
upgrade_required |
a plain HTTP request, not a WebSocket upgrade |
| 429 |
too_many_sessions |
no Osprey Live session can open right now: this account already has as many open as it may, or the server's sessions are all in use. Retry shortly |
| 429 |
hearing_busy |
live hearing capacity is full right now. Retry shortly |
| 503 |
hearing_unavailable / service_unavailable |
a part of the agent is not serving right now. Retry shortly |
POST /v1/video/summaries — Lark-V
MiniCrow-specific. OpenAI has no shape for "video → timeline", so this one is ours.
curl https://api.minicrow.com/v1/video/summaries \
-H "Authorization: Bearer mc_..." -F file=@clip.mp4 -F model=lark-v-large
{"model":"lark-v-large",
"summary":"A colorful test pattern is shown while a voiceover announces a meeting in Pune.",
"timeline":[{"t":0,"text":"A test pattern displays while a voice says, \"Kal shaam 5:00 baje…\""}],
"usage":{"prompt_tokens":288,"completion_tokens":183,"cost":9.5278,"cost_currency":"INR_paise"}}
⚠️ lark-v-large is the most accurate tier. It sees the picture and hears the sound together: speech in the
clip appears in the timeline, in the script it was spoken in.
⚠️ Billed on what the clip used, not per second — a per-second rate would be a guess dressed as a price. Every
response carries its cost.
lark-v-nano and lark-v-mini
These two return the same summary and timeline, and the words spoken with their times too. lark-v-mini is the more
accurate of the two. lark-v-nano is coming soon: until it opens it answers 503 tier_not_deployed, and the
message names the tier to use. Everything below applies to both:
curl https://api.minicrow.com/v1/video/summaries \
-H "Authorization: Bearer mc_..." -F file=@clip.mp4 -F model=lark-v-mini -F effort=high
{"model":"lark-v-mini",
"summary":"Vertical colour bars with a moving diagonal stripe, while a voice repeats a meeting reminder.",
"timeline":[{"t":0,"text":"Colour-bar test pattern; a voice says \"Kal shaam 5:00 baje…\""}],
"x_minicrow":{"tier":"lark-v-mini","effort":"high","heard":true,"cost_known":true},
"usage":{"prompt_tokens":2058,"completion_tokens":1104,"seconds":24,
"cost":18.983,"cost_currency":"INR_paise"}}
| field |
meaning |
x_minicrow.heard |
whether the clip had a speech track and it was transcribed |
transcript |
the speech track, as heard |
transcript_segments |
the same transcript a sentence at a time, each with when it is said — [{"start":15.04,"end":20.0,"text":"…"}], in seconds; estimated, see below. Absent when nothing in the track was voiced |
usage.seconds |
seconds of speech in the clip; 0 for a silent clip |
⚠️ A silent clip is charged nothing for speech and reports heard: false.
⚠️ The whole clip is summarised, however long — never only its first minute.
⚠️ At most 10 minutes of video on these two tiers — 413 video_too_long beyond that.
⚠️ lark-v-nano is the lighter tier, not a cheaper picture. Its price is close to lark-v-mini's, and on
accuracy lark-v-mini is currently the stronger of the two. Both tiers are announced separately in
GET /v1/models: one of them serving never implies the other does.
⚠️ What is said sits in the timeline where it is said (since 0.62.0), and transcript_segments carries each
sentence with its time. The times are an estimate: on one voice without music they land within about a
second.
⚠️ If the model does not return parseable JSON you still get its answer, as raw, with
timeline_parsed: false. The model has already done the work; throwing the response away would charge you for
nothing.
Room to write the answer: output_budget and answer_cut_off
Every Lark-V response reports x_minicrow.output_budget — the most output tokens the answer could use — and
x_minicrow.answer_cut_off, true when the answer ran out of it before finishing. Lark-V reasons before it
writes, and the reasoning is paid out of the same budget: a 60-second Hindi clip in script=native at effort=high
spent 4,094 of its 4,096 tokens thinking and returned an empty answer. Since 0.62.0 the budget is the preset's —
2,048 (mid), 4,096 (high), 8,192 (max) — multiplied by 3 when the summary or the quoted speech is
written in an Indian script, by 2 in another script that is not Latin (Arabic, Russian, Japanese, Korean), and
by up to 3 more on longer clips on lark-v-nano and lark-v-mini, capped at 32,768. The same
60-second clip now finishes in 4,552 tokens. It is a ceiling, not a charge: you pay for the tokens the answer used.
An answer that is still cut off comes back with what it wrote, answer_cut_off: true, and — if the JSON did not
survive — raw; script=latin or a lower effort needs less room.
language and script
-F language=ta -F script=native
Declare the language people speak in the video, as on /v1/audio/transcriptions. For the ready languages —
mr hi ta te bn gu kn ml en es de fr it ar — it changes four things:
- How speech is quoted. Every tier knows which language is spoken and quotes it in that language's own
script (
script=native, the default), romanised (latin), or with no alphabet asked for (auto), never translated.
- Which lane hears the speech on
lark-v-nano and lark-v-mini: the same lane an audio file in that language
gets on the matching Lark tier, billed at that lane's rate. x_minicrow.speech_lane says which.
- How the speech track is transcribed: as
/v1/audio/transcriptions transcribes that language, treating the
audio as a video soundtrack — only spoken or sung words are written.
- The language of the answer. The
summary and every timeline text (the descriptions of what is seen) are
written in the declared language — in its own script for script=native, romanised for latin, in English for
en. What people say is still quoted as the first point says; summary_language below picks another language.
x_minicrow.language echoes the language when it was used, and is "" when none was declared or the language is not
one of the above — those requests are served exactly as before. An unknown script is 400 unknown_script on every
tier.
summary_language
-F language=ta -F summary_language=en
Choose the language of the summary and the timeline descriptions yourself — a Tamil video summarised in English,
as above. It overrides the default from language, takes the same script, and accepts any of the 21 live
languages: mr hi bn te ta gu kn or ml pa as en es fr de pt it ru ar ja ko
(a region such as en-US is ignored). It works on every tier and never changes how speech is heard or quoted.
- With
language too, speech follows language and the answer follows summary_language.
- On its own, only the answer's language changes; everything else is served as it is with no
language. The
code-mixed-Hindi rule then applies to quoted speech only, so a hi or mr summary can be written in Devanagari.
- How faithfully each tier follows a
summary_language different from language has not been measured yet.
- With neither, no answer language is set, exactly as before.
x_minicrow.summary_language echoes the language the answer was asked to be written in ("" when none was chosen).
Anything outside the list is 400 unknown_summary_language, before anything is decoded or billed.
⚠️ Not measured yet on real videos: how faithfully each language is followed.
effort — mid · high · max
-F effort=max
| preset |
what you get |
output budget |
mid |
a handful of entries — the key moments |
2,048 tokens |
high (default) |
an entry every few seconds |
4,096 |
max |
an entry for every distinct moment |
8,192 |
The response echoes the preset in x_minicrow.effort.
An unknown preset is a 400, not a silent default — a caller who typed maximum asked for something specific.
At most 20 MB — a longer video is a job queue, which this
endpoint is not. Over that is 413 file_too_large.
Errors specific to this endpoint
| status |
code |
when |
| 400 |
unsupported_video |
the bytes are not a video container, or the container carries no picture |
| 400 |
empty_video |
it is a video and it has no measurable length — nothing to summarise |
| 400 |
unknown_effort |
a preset that does not exist; nothing was decoded and nothing was billed |
| 400 |
unknown_script |
script is not native, latin or auto; nothing was decoded and nothing was billed |
| 400 |
unknown_summary_language |
summary_language is not one of the 21 live languages; nothing was decoded and nothing was billed |
| 413 |
file_too_large |
over 20 MB |
| 413 |
video_too_long |
over 10 minutes, on lark-v-mini |
| 503 |
tier_not_deployed |
a tier this deployment does not serve; the message names the one to use instead |
GET /v1/models
Public — no key. A client has to be able to see what it may ask for, and the prices are the product.
{"object":"list","data":[
{"id":"osprey-flash","object":"model","owned_by":"minicrow","label":"Osprey Flash",
"modes":["speed","intelligence","max","auto"],"supports_tools":true,"available":true},
{"id":"osprey-flash-lite","object":"model","owned_by":"minicrow","label":"Osprey Flash Lite",
"modes":["speed","intelligence","auto"],"supports_tools":false,"available":false,
"unavailable":{"code":"coming_soon",
"reason":"Osprey Flash Lite is coming soon. Use osprey-flash or osprey-pro."}}
]}
⚠️ available: false means do not send a request — every call to that id answers 503 tier_not_deployed.
unavailable.code is coming_soon for a model that is announced and not open yet, and tier_not_deployed for one
that is not serving; reason names what to use instead. Filter on the field rather than on the label — the list
changes as tiers open.
⚠️ They are listed rather than hidden, on purpose. A product that is in the price list and absent from the
catalogue is one you meet as unknown_model — which reads as a mistake in your code rather than a gap in ours.
osprey-live joined the list when access opened to every account (0.46.0). Its entry carries a languages list,
each with a status of ready (today hi and en) or coming_soon; its section
is the reference.
The flag is derived from the same function the handlers refuse with, so what this endpoint advertises and what a
call actually does cannot drift apart.
GET /healthz
Public. {"ok":true,"service":"minicrow","version":"0.1.4"}. Liveness only — it deliberately does not touch
the database, because a probe that did would restart every pod at once on one blip.
/readyz exists but is not published: it reports whether Postgres and Redis are reachable, which is
reconnaissance to a stranger and useless to a customer.
Errors
| HTTP |
code |
When |
| 400 |
invalid_json |
the body is not JSON |
| 400 |
model_required |
no model field |
| 400 |
unknown_model |
no such tier — the message lists the ones that exist |
| 400 |
unsupported_audio |
the recording's container is not one MiniCrow can measure, and the charge is per second of it |
| 400 |
empty_audio |
the recording is valid and zero seconds long; a model handed silence writes words nobody said |
| 400 |
script_mismatch |
an English voice was asked to read Devanagari |
| 400 |
emotion_not_supported |
an emotion was sent to a lane that has none |
| 400 |
lane_not_supported |
a voice you created was asked for on the expressive lane, which has its own fixed cast |
| 400 |
attestation_required |
a clone was sent without attest=true |
| 400 |
no_reference_clip |
a multipart create with no file |
| 400 |
unusable_clip |
the recording could not be decoded, or is not speech we can use |
| 400 |
clip_duration |
the recording is shorter than 3 seconds or longer than 30 |
| 400 |
description_required |
a JSON create with no description |
| 404 |
unknown_voice |
no such voice on this account — the same answer as a typo, on purpose |
| 409 |
voice_limit_reached |
this account already holds the maximum number of voices |
| 413 |
clip_too_large |
the upload is far larger than a reference clip |
| 429 |
voice_rate_limited |
this key has made too many voices in the last hour or day |
| 503 |
voices_unavailable |
this deployment serves the built-in voices only |
| 503 |
tier_not_deployed |
a tier that is not serving on this deployment |
| 400 |
unknown_mode |
the tier exists, the suffix does not. Distinct from unknown_model on purpose: a client switches on the code, and one of those means "stop using this model" while the other means "fix one word" |
| 401 |
invalid_api_key |
missing, malformed, unknown or wrong key — always the same answer |
| 401 |
key_revoked |
the key was revoked |
| 401 |
key_expired |
the key reached the expiry it was created with. Distinct from key_revoked on purpose: nobody stopped this key, its date passed — the fix is to mint a new one, not to find out who revoked it. A key created without an expiry never returns this |
| 402 |
insufficient_credit |
prepaid balance at or below zero |
| 429 |
— |
rate limited at the edge; see below |
| 503 |
model_unavailable |
the model could not be reached. Recorded but not charged — a call that did not happen is not billed. Carries X-MiniCrow-Origin-Status: 502 |
| 400 |
model_refused_request |
the model refused the request as sent. Check the parameters and the message shape |
| 400 |
context_too_long |
the prompt plus max_tokens is longer than the lane can take |
| 429 |
rate_limited |
the lane is busy. Try again shortly, or use another mode |
| 503 |
model_timeout |
no answer in time. Carries X-MiniCrow-Origin-Status: 504 |
| 503 |
store_unavailable |
the key store could not be reached |
| 503 |
embeddings_failed, speech_failed, transcription_failed, video_failed |
that service could not answer. Carries X-MiniCrow-Origin-Status: 502 |
| 404 |
unknown_endpoint |
no such path. The message lists the endpoints that exist |
| 405 |
method_not_allowed |
the path exists and does not take that method |
| 413 |
request_too_large |
the body is past the gateway limit; audio is capped at 25 MB |
| 500 |
internal_error |
the gateway itself failed. Nothing was charged |
| 426 |
upgrade_required |
a plain HTTP request to the live endpoint, which only speaks WebSocket. The live endpoint's other refusals are in its own table |
Why a gateway failure is a 503 and not a 502
⚠️ Nothing from this API ever answers 502, 504 or 52x — and that is not a taste in status codes. Cloudflare
sits in front of api.minicrow.com, and it mints those statuses itself for its own edge→origin failures. It
cannot tell ours from its own, so it discards our body and serves 16 bytes of error code: 502 in
text/plain. Measured on the live host, one status at a time:
| origin answered |
what reached the customer |
| 200, 400, 429, 500, 501, 503, 505, 507, 508, 509, 527, 530 |
passed through byte-for-byte |
| 502, 504, 520, 521, 522, 523, 524, 525, 526 |
body discarded, replaced with error code: NNN, text/plain |
The zone setting that would stop it — origin_error_page_pass_thru — is Cloudflare Enterprise-only and reads
{"value":"off","editable":false} here. So the fix is on our side: a status the edge would swallow never
leaves this API. What was a 502 or a 504 is answered as 503 Service Unavailable, with:
- the same envelope and the same
error.code — the field you branch on is unchanged,
Retry-After: 5, which a 502 never carried, and
X-MiniCrow-Origin-Status, naming the status the gateway itself decided, so a demoted 504 stays
distinguishable from a 503 that always was one.
⚠️ Every failure is JSON, including the ones no handler wrote. A mistyped path, a wrong method, a body over
the limit and an internal panic all come back as the envelope above. There is no response from this API that a
client parsing error.code cannot parse.
Rate limit
50 requests/second sustained, burst 100, per client IP (keyed on Cf-Connecting-Ip). This is far above what
a real caller generates — a chat call takes seconds — because it does not exist to ration the product. It exists
to bound the cost of refusing: every wrong key costs one argon2id hash at 19 MiB before it can be rejected.
What a failure will not tell you
⚠️ A failure is described in MiniCrow's words. The message and the code say what you can act on —
context_too_long and rate_limited mean exactly what they say.
⚠️ A 401, 402 or 404 is always about your own key, account or request. A lane that fails for a reason of ours
is answered as model_unavailable. If it came back as 401 you would check your key; as 402 you would top up.
Neither would help: those are our problems with a lane, not yours with your account.
⚠️ Codes changed on 2026-09-11. Four earlier 5xx codes were replaced by the table above. Done while the only
client was one we control; a code is a contract, and waiting would have cost more.
What is not here yet
|
|
| Lark Live |
preview, enabled per account — every other account is refused 503 tier_not_deployed |
| Osprey Flash Lite |
coming soon — refused 503 tier_not_deployed; the message names osprey-flash and osprey-pro |
| Lark Nano |
coming soon — refused 503 tier_not_deployed; the message names lark-mini |
| Lark-V Nano |
coming soon — refused 503 tier_not_deployed; the message names lark-v-mini |
Everything else in the catalogue serves. GET /v1/models is the live answer; this table is a summary of it and can
only ever be staler.
When a mode is unavailable but the model is not
GET /v1/models lists the modes a model accepts, and names any that refuse. The example below is the shape; today
every listed mode serves, lark-large:max included:
{ "id": "lark-large", "available": true, "modes": ["max"],
"unavailable_modes": [{ "mode": "max", "code": "tier_not_deployed",
"reason": "This mode is announced but not wired up yet. Leave the mode off and Lark Large will serve." }] }
⚠️ A TIER THAT SERVES CAN CARRY A MODE THAT DOES NOT, AND THE CATALOGUE USED TO SAY NOTHING. lark-large
answers and lark-large:max returned 503, so a customer read available: true and got a refusal from a mode
this very response advertised. Found by calling every mode of every listed model, not by reading the list.
⚠️ lark-large TAKES ONE MODE, max, AND NOTHING ELSE. default and indic are chosen for you from
the language you declare — asking for them is 400 unknown_mode, which is correct: they are how MiniCrow
describes the lane it picked, not something to request.