Back to the guide

< Full API reference >

Every endpoint, field and error

The complete reference, generated from the same file the API is built against. For a shorter path in, start with the guide.

MiniCrow for AI agents

This file is for the agent (or engineer) integrating MiniCrow into an existing system. It says what each model is, when to use it, the exact request shape, the rules that decide accuracy, and what to avoid. Everything here is measured on MiniCrow's own benchmarks; where a number is directional it says so. The full reference is /docs (the API reference); this page is the short, prescriptive version. Base URL https://api.minicrow.com, header Authorization: Bearer <key>, OpenAI-compatible wire format. Every response carries the price of the call in paise.

The models, in one table

Model What it does Call it when Do not call it when
osprey-flash (:speed · :intelligence · :max · :auto) Text chat and tool calling, OpenAI-compatible You need an LLM answer, a JSON extraction, a tool call from text You have audio (use Lark) or a live phone call (use Osprey Live)
osprey-pro (:speed · :intelligence · :max · :auto) The same, larger model Long reasoning, hard extraction, code Simple replies — osprey-flash is cheaper and faster
osprey-flash-lite — coming soon The smallest Osprey Not yet: every call is refused 503 tier_not_deployed. Use osprey-flash:speed Until GET /v1/models lists it available: true
lark-nano · lark-mini · lark-large Speech to text from a recording, 21 languages, code-mixed. Nano is coming soon, mini the default, large the most accurate You have a finished recording (call, voice note, meeting) You need words while the caller is still speaking (Lark Live)
lark-mini live (/v1/audio/transcriptions/live) Speech to text over a WebSocket, turn by turn, ~0.7 s after the caller stops A live call where your own system decides what to say You want MiniCrow to answer the caller (Osprey Live)
osprey-live (/v1/agent/live) A complete voice agent over one WebSocket: hearing → your instructions and tools → speech. It also sees: send pictures from the caller's camera beside the audio A phone agent that books, checks, sends, hands over; an assistant on glasses, a robot or a kiosk that answers about what its camera shows Batch work, or languages other than Hindi and English (coming)
pica-nano · pica-small · pica-large Text to speech, Hindi and English, streaming, telephone format; large is the most natural You already have the words and need audio —
lark-v-nano · lark-v-mini · lark-v-large Video → summary and timeline JSON (what is said and seen, by second). Large is the most accurate: it sees and hears together; nano is coming soon A recording with a picture that matters Audio only (Lark is cheaper)
minicrow-embed (/v1/embeddings) · minicrow-rerank (/v1/rerank) Embeddings (dense and sparse) and reranking, for search and RAG You index or search your own text —

Every generative model above answers as MiniCrow's: osprey-flash, osprey-pro, or your Osprey Live agent's persona.

Rules that hold everywhere

  1. Declare the language. language on every speech call (hi, mr, en, …). MiniCrow never guesses; the declaration picks the lane and tunes the transcription to that language. A wrong or missing language costs more accuracy than anything else.
  2. Declare the alphabet for Indian languages. script=latin (romanised, as people type in chat) or script=native. The transcript follows it; English words inside a sentence stay in English either way.
  3. Telephony audio is fine as it is. 8 kHz μ-law/A-law or 16-bit PCM straight from the trunk. Do not upsample or "enhance" it; MiniCrow's speech lanes were measured on 8 kHz calls.
  4. One speaker per channel. For calls, send each party on its own channel or its own session. Mixed audio makes the agent hear itself, and no model separates two people in one channel well.
  5. Hints are optional and small. Accuracy is built to hold with none. Send the names this recording or call will contain (vocabulary), not a glossary of everything you sell; a long list costs more on every turn and can add words that were never said.
  6. Read the response. x_minicrow says which lane served, why, and the exact price. Log it; it is your bill and your audit trail.

Lark (recorded speech → text): best practice

Endpoint POST /v1/audio/transcriptions, multipart: file, model=lark-mini, language, script, optional domain, vocabulary, abbreviations, prompt, diarize.

  • Send WAV/OGG-Opus/MP3/M4A/WebM. The bill is per second, read from the file header.
  • Two-channel calls: diarize=true with a stereo file labels the parties from the channels — the cheapest and most accurate speaker labelling MiniCrow offers. One channel: diarize=true runs the speaker pass (slower, ₹ per hour).
  • The hint fields, in order of value: vocabulary (names spelled as you want them), abbreviations (SHORT = long), domain (one line, who talks to whom), prompt (anything else, background only). Measured: eight names and a one-line domain moved accuracy within the margin on read speech and helped only where the names were actually spoken.
  • Never paste your agent script, FAQ or a previous transcript into prompt: on a near-silent clip Lark can return it as the transcript. MiniCrow refuses a transcript that is a copy of one of Lark's own reference lines; it cannot know yours.
  • Long recordings: lark-mini has no duration cap below 25 MB. lark-nano is coming soon.
  • A WAV that certainly holds no speech — silence, a steady tone or hum, line noise, a busy or ringing tone — comes back text: "" with x_minicrow.no_speech: true, not billed. On a call split by channel, a channel with no speech comes back as an empty segment with no_speech: true. A tone that switches on and off can still be transcribed.
  • Two channels: channel 1 is the left one. A split call returns one segment per channel, each with its own script_honoured and script_repaired; the call is honoured only when every channel is.
  • A transcript that loops is never returned: it is transcribed again without your hints, or the loop is kept once; x_minicrow.repetition says which (retried_without_hints or collapsed).
  • verify_vocabulary=true checks every vocabulary name the transcript writes against a second transcription made without the vocabulary, and replaces a name that was not heard with what was. It costs one more transcription of the recording, only when a name was written; x_minicrow.vocabulary_check lists what was confirmed, removed or left.

Lark Live (streaming speech → text): best practice

Endpoint GET /v1/audio/transcriptions/live (WebSocket). Query: model=lark-mini, language, script (required for Indian languages), encoding (linear16|mulaw|alaw), sample_rate (8000|16000), endpointing_ms (300–1000), vad (server|client), domain, vocabulary, context (repeatable).

  • Send audio at real-time pace in 20 ms–1 s frames. Read transcript.partial (stage draft, ~0.7 s after the caller stops) to act fast and transcript.final (~1.5 s later) to store; the final corrects the draft.
  • endpointing_ms is the trade: 300 answers sooner and splits sentences; 600–800 keeps a slow speaker's sentence whole. Default 400. Use vad=client only if your telephony stack already detects turns.
  • vocabulary (up to 40 terms) is used on every turn AND is boosted inside the draft for hi, mr and en — the languages the boost is measured on; the other 18 use it as a spelling hint only, without the boost, until measured. List names, never everyday words; for hi/mr add the Devanagari spelling as its own term (Orbit, ऑर्बिट). Measured: Hindi/Marathi boosted names caught 5 of 7 against 2 of 7 without, word error unchanged; English 13 of 13 against 12, word error 14.3 → 12.4.
  • hi and mr drafts are locked to Devanagari (English words in English letters stay), vocabulary or not.
  • Reconnect mid-call by opening a new session with the last finals in context.
  • Speaker labels are not offered live; one session per channel is the way.

Osprey Live (voice agent): best practice

Endpoint GET /v1/agent/live (WebSocket). Query: language (hi|en), voice, encoding, sample_rate, output (mulaw_8k for a phone line), endpointing_ms, barge_in, transcripts. Then one session.configure message with your agent spec, then audio in, audio + events out. You run every tool; MiniCrow sends tool.call and waits for your tool.result.

The spec decides whether tools get called. Measured on 99 turns where a tool was required (invented clinic, car service, courier and property calls, Hindi and English, plus real Marathi calls): the default brain made 59 % of them; with the caller's words heard perfectly the same brain made 81 %, and the larger brain 89 %. Hearing, then spec wording, are the levers you own.

The spec, field by field, and the rule for each

Field Rule
brief Who the agent is, how it speaks (language, script, one short sentence, numbers as digits), the facts it may state, the goal of the call in order. Say "call the tool for anything not stated here". Keep it short: it is sent every turn
tools[].when_to_call Start with "CALL THIS", name situations not test sentences, say BEFORE ("CALL THIS BEFORE saying any time"), close the escape ("never answer from memory")
tools[].when_not_to_call For actions: "already done on this call", "number still being dictated"
tools[].parameters Require only what the tool cannot run without; ask optional details after the result. A lookup or booking with more than 3 required fields waits for all of them. Describe a phone field as a value ("caller_id when they say to use this number, otherwise the number they give"), never as steps ("dictated and confirmed")
tools[].kind lookup (reads), booking (checks, commits nothing), action_needs_confirmation (the caller will notice), other
tools[].confirm: "runtime" Put it on every action whose arguments the caller dictates (phone, name, email). The first call is read back digit by digit and a misheard value is asked again; the call reaches you only after a yes
tools[].acknowledgement 3–5 spoken words that promise nothing ("Ji, time dekh leti hoon.")
examples One invented example per tool you most need called, one where the lookup does not answer, one number dictation. Everything in them invented; never a real number
hearing.vocabulary Names your callers say: people, products, places, codes, up to 100, with spoken Devanagari forms for Hindi and Marathi callers ({"term":"Orbit","spoken":["ऑर्बिट"]}). Names, never everyday words. Tool enum values are boosted on their own. Boosted in the hearing for hi, mr, en; Hindi and Marathi hearing is locked to Devanagari
speech.reprompt, greeting_words, never_claim_before_result, block_before_result What to say on an unreadable turn; words that mean a greeting; outcome words the agent may not say before a tool result
history The greeting you already played, or the call so far on a reconnect

session.configured.warnings checks your spec against these rules and names the field. Fix them before going live.

The camera (vision). Send a picture from the caller's camera every 1–3 seconds as {"type":"input.image","data":"<base64 JPEG, PNG or WebP>"}, after session.configured (which carries vision on a session that takes pictures). At most 192 KB a picture — 640 to 1024 pixels wide is plenty; at most one a second is kept. When the caller finishes a turn, the reply is written with up to 3 of the pictures taken while they spoke, so "what is this?" is answered about what the camera showed. Only the turn's own pictures are seen; what the agent said about them stays in the history as text. Pictures are billed as brain input tokens — about 1,100 each.

Runtime you can rely on (whatever the prompt says): no outcome claim before a tool result; no second greeting; an identical successful action is not repeated; at most 3 brain passes and 2 tool calls per caller turn; asked what it is or who made it, it answers only as your persona; a tool call the agent writes into its words instead of making it is never spoken; runtime confirmation as above — and a yes to the read-back releases the held call on the next reply; a phone number the caller never said — a placeholder such as 9876543210, or one copied from your examples — is never sent to a tool: the agent is told to use caller_id (the caller said "use this number") or to ask for the number.

Latency shape: first audio ≈ 2.3–2.6 s after the caller stops on a cache read; the first 3–4 turns of a call are warmed at connect. A tool turn speaks the acknowledgement, calls the tool, and answers from the result. Play audio as it arrives; do not buffer.

Choosing the brain: the default brain is the cheaper one (our measured 59 % → 66 % with hearing vocabulary on the turns above); the larger brain scores 74–76 % at about 1.5× the brain price. Ask MiniCrow to switch your account if required tool calls matter more than the difference.

Osprey (text chat): best practice

POST /v1/chat/completions, model = osprey-flash | osprey-flash:speed | osprey-flash:intelligence | osprey-flash:max | osprey-flash:auto | osprey-pro (same modes). osprey-flash-lite is coming soon.

  • :auto chooses the lane per request from the request itself (tools, length, script) and may step up once after a verifiable failure; both attempts are charged. Name a mode when you need a fixed price.
  • Send X-MiniCrow-Session with a stable id per conversation so MiniCrow keeps the same lane across turns.
  • Tools: OpenAI tools + tool_choice. Keep schemas tight (enum where the values are known); the response reports tool_calls_valid.
  • Streaming: stream: true; the final chunk carries usage and the cost.
  • Thinking counts toward max_tokens and is billed; if it uses the whole budget the answer is empty with finish_reason: "length". Give thinking modes 2,000+ tokens. On :speed, reasoning_effort: "none" switches thinking off for that request: about 3 s instead of 8 and never empty — good for summaries and extraction, worse for calculations. :intelligence always thinks.
  • The model presents itself as osprey-… by MiniCrow.

Embeddings and reranking (search and RAG): best practice

  • POST /v1/embeddings with model=minicrow-embed and input (a string or up to 256 strings): each text gets a 1,024-number dense vector and sparse lexical weights in the same response. Keep both — hybrid search (dense + sparse) finds the exact names and codes that dense vectors alone miss. Send text, never token-id arrays.
  • POST /v1/rerank with model=minicrow-rerank, query, documents, top_n (Cohere's shape): rerank your top 20–50 candidates from search, not the whole collection. Billed on the query and the documents.

Pica (text → speech): best practice

POST /v1/audio/speech with model, voice, input, response_format (mulaw_8k for a phone line), stream: true for first audio as soon as it is ready.

  • Pick the tier: pica-nano (₹21/h) for the eleven built-in voices, your own cloned voices, emotions and first_clause/chunk; pica-small (₹32/h) for its 25 Hindi and English voices; pica-large (₹60/h) when the voice has to sound the most natural. List a tier's voices with GET /v1/audio/voices?model=….
  • A built-in or cloned voice is always spoken by pica-nano, whatever model says; X-Model tells you which tier answered. lane, emotion, chunk and first_clause on small/large are a 400, not ignored.
  • 429 tier_busy on small/large: retry shortly, or fall back to pica-nano.
  • Write the text as it should be said: digits for numbers, Devanagari for a Hindi voice. The bill is per second of audio received, so shorter wording costs less.
  • Voice cloning (POST /v1/audio/voices) needs the speaker's consent; the reference explains what is stored.

Lark-V (video → timeline): best practice

POST /v1/video/summaries with the file, model (lark-v-mini | lark-v-large; lark-v-nano is coming soon), language, script, summary_language, effort (mid|high|max). The answer is JSON: a summary and a timeline of {t, text} in seconds. lark-v-large is the most accurate — it sees the picture and hears the sound together. Use effort=mid unless frames matter. Billed on what the clip used; every response carries its cost.

Integration checklist

  • Base URL and key set; x_minicrow logged from every response.
  • language (and script for Indian languages) declared on every speech call.
  • Telephony audio sent as-is (8 kHz, one party per channel).
  • Live: partials used to act, finals stored; endpointing_ms tuned on your own calls.
  • Osprey Live spec: session.configured.warnings empty; confirm: "runtime" on dictated-value actions; hearing.vocabulary filled with names (Devanagari spoken for hi/mr); examples invented.
  • Tools answered within tool_timeout_ms; results short JSON; call_id echoed.
  • Reconnect paths implemented (close 1012/1013 → new session with history / context).
  • Osprey Live with a camera: pictures sent after session.configured, one every 1–3 s, each under 192 KB.

Prices (customer, list; ₹)

  • Lark, per hour of audio: nano ₹10, mini ₹20 (batch and live), large ₹30 — ₹45 when an Indian language is declared, ₹60 for lark-large:max.
  • Osprey Live: hearing ₹20 an hour + brain per token by type (≈ ₹13.5 per call-hour on the default brain) + voice ₹9 per 10,000 characters (≈ ₹32 per call-hour) ≈ ₹66 per call-hour all-in. Camera pictures add brain input tokens: about 9 paise on a turn with three.
  • Pica, per hour of audio: nano ₹21, small ₹32, large ₹60.
  • Osprey chat, embeddings and rerank: per million tokens — GET /v1/models lists every rate.
  • Lark-V: billed on what each clip used.

GET /v1/models and the price in every response are the authority.


MiniCrow API

Base URL https://api.minicrow.com Compatibility OpenAI Chat Completions. An existing OpenAI client works by changing two things: the base URL and the key.

Everything on this page is live and was verified against the deployed service, not against a mock.


For AI agents and quick integration: /agent.md is the short, prescriptive version of this reference — every model, when to call it, the request shape, and the rules that decide accuracy (docs/AGENT.md in the repo).

Authentication

Authorization: Bearer mc_<prefix>_<secret>

A key looks like mc_96bf0550e045_kQ3v…. The middle segment is the prefix — it identifies the key in a listing or a log line and is not secret. The whole string is the secret and is shown once, at creation, and never again: only an argon2id hash of it is stored, so a lost key is replaced, not recovered.

Every failure to authenticate answers the same 401 invalid_api_key with the same message, whether the key does not exist or the secret is wrong. A different answer for each would tell an attacker which half to keep guessing.

Credits and keys

Credits belong to the account. Keys draw on them. Make as many keys as you like — one per service, one per environment — and they all spend the same balance. A new key works immediately; there is nothing to top up on it.

A key may carry its own spend limit. That is what a key for a script, a contractor or a staging environment is for: it stops at its own cap and the rest of your balance stays usable.

402 insufficient_credit the account is out of credit
402 key_limit_reached this key has hit its own limit; the account has not
403 account_suspended the account is suspended. Not a credential problem — rotating the key will not help

Both are checked before the model is called, so neither costs you anything.

⚠️ A key with no limit set has no limit — it is not a limit of zero. A key created without one works.

⚠️ A key's limit is a LIFETIME limit. Nothing resets it: spent counts everything that key has ever spent, so a cap of ₹500 is ₹500 for the life of the key, not ₹500 a month. There is no billing period in the service yet. Raise the cap, or mint a new key, when it is reached.


POST /v1/chat/completions

Standard OpenAI request body. Two things differ, and both are additions rather than changes.

Every MiniCrow model presents itself as MiniCrow's

0.42.0. A generative model here answers as the product it was called as — osprey-flash, osprey-pro, or the persona your Osprey Live spec gives it — and as MiniCrow's. Transcription and summarising models (Lark, Lark Live, Lark-V) produce only a transcript or a timeline and are not asked such questions.

The model id carries the mode

You send You get
osprey-flash the tier's default lane. Reported as default:<mode>, not as a request you made.
osprey-flash:speed that lane, always. An explicit mode is a hard choice, not a hint.
osprey-flash:intelligence "
osprey-flash:max "
osprey-flash:auto MiniCrow picks the mode for each request, from the request itself (tools, length, script).

GET /v1/models lists every tier and the modes it offers: what you buy is the lane.

Every response says which lane answered and why

"x_minicrow": {
  "requested_mode": "auto",
  "served_mode":    "intelligence",
  "route_reason":   "S:tools",
  "cost_known":     true
}

This is deliberate and is not going away: a choice that is hidden cannot be debugged, so the lane MiniCrow picked and the rule that picked it are in every response.

⚠️ The lane is the product. You buy a lane at a price; the lane is what you can act on.

route_reason distinguishes explicit:<mode> (you named it), default:<mode> (you sent a bare model id) and MiniCrow's own reasons for the mode it picked, such as S:tools.

Cost is in the response, in paise

"usage": {
  "prompt_tokens": 88, "completion_tokens": 60,
  "cost": 0.6458,
  "cost_currency": "INR_paise"
}

⚠️ usage.cost is always what MiniCrow charges you, and it is the same number in a streamed response as in a non-streamed one. A typical short call costs a fraction of a paisa, which is why the field is decimal and why balances are held to six decimal places — rounding each call to a whole paisa would report a busy month as free.

cost_known: false means the call's cost could not be measured and the lane has no flat rate, so nothing was charged. That is a gap you should see, not a discount.

Thinking, and the room it needs

Some modes think before they answer. Thinking tokens count toward max_tokens and are billed as output; the number is in usage.completion_tokens_details.reasoning_tokens. If the budget runs out while the model is still thinking, the answer comes back empty with finish_reason: "length" — and is billed. Give a thinking mode room (2,000 tokens or more for a long input), or switch thinking off where the task does not need it.

Mode Thinks reasoning_effort
speed yes, by default low · high · max set how much; none (or minimal, or "reasoning": {"enabled": false}) switches it off for this request
intelligence always low · high · max set how much; it cannot be switched off
max yes minimal · low · medium · high · max

Any other value is ignored. Measured on osprey-flash:speed over 60 requests: with thinking off it answered in 3.0 s against 7.9 s (median), never came back empty (5 of 60 did with thinking at max_tokens 2,000, and 23 of 60 would have at 400), returned valid JSON 30 of 30 times against 24, and was judged better on call summaries — but worse on classifications, chat replies and rewrites, and got 5 of 10 calculations wrong. Switch it off for summaries and extraction; keep it for anything that has to be worked out.

Streaming

"stream": true returns text/event-stream, terminated by data: [DONE].

Every chunk carries the model id you called, and the final usage chunk carries usage.cost in paise plus x_minicrow. Nothing else about how the lane was chosen is included. Cloudflare does not buffer these streams; chunks were measured arriving 20–30 ms apart through the public host.

Each event is written and flushed on its own, the moment it arrives: one data: {…}\n\n per event. A failure that happens before the first event — a lane that is down, a rate limit, a model that accepted the call and then said nothing — is not a stream at all: it is the ordinary JSON error envelope with its own status, exactly as without stream, and it is not charged.

A stream that breaks after it started ends on an error event, in the same envelope as every other error:

data: {"error":{"message":"The answer was cut off before it was complete. Try again.","type":"api_error","code":"model_unavailable"}}

It is the last event, and there is no data: [DONE] after it — so an OpenAI client raises instead of reading the partial answer as a whole one. When the model itself reports a failure mid-answer, that chunk keeps its place with an error object in the same shape (model_unavailable, rate_limited, model_timeout, …).

A stream that breaks part-way is still charged for what was produced: a partial answer you received is not free to make. Hanging up does not stop the charge either. The cost arrives in the final usage chunk and covers the whole generation, so when you disconnect MiniCrow stops sending but lets the generation finish and charges what it used.


POST /v1/embeddings

minicrow-embed returns dense and sparse vectors in one call. The sparse half (lexical weights) is what hybrid retrieval needs, and most embedding APIs do not return it.

{"model":"minicrow-embed","input":["kal meeting hai Pune me","tomorrow there is a meeting"]}
{"object":"list","model":"minicrow-embed","data":[
  {"object":"embedding","index":0,"embedding":[…1024 floats…],
   "sparse":{"indices":[5,9,…],"values":[0.7,0.2,…]}}],
 "usage":{"prompt_tokens":12,"cost":0.0025,"cost_currency":"INR_paise","prompt_tokens_estimated":true}}

⚠️ prompt_tokens_estimated: true is not decoration. Embedding produces no token count of its own, so the count is MiniCrow's estimate, made with the same per-script estimator MiniCrow uses when it picks a chat mode. Every other endpoint bills from a measured figure; this one says when it cannot.

at most 256 inputs capacity is shared with live retrieval; an unbounded batch stalls everybody's search
pre-tokenised input is refused OpenAI accepts token-id arrays. MiniCrow's tokenizer is not OpenAI's, so those ids mean nothing to it — embedding them would return confident vectors for the wrong text
a short batch is a 503 embeddings_failed the results are positional. Three vectors for four texts would put every vector on the wrong text, with no error anywhere. It is decided as a 502 and answered as a 503 — see Errors

POST /v1/rerank

Cohere's shape, because OpenAI has no rerank endpoint and every client that speaks rerank speaks this one.

{"model":"minicrow-rerank","query":"Pune meeting kab hai",
 "documents":["Delhi ka flight subah 6 baje","Pune me meeting kal 3 baje hai"],"top_n":2}

⚠️ Billed on query + documents, not documents alone. A cross-encoder reads the query once per document, so counting only the documents would under-count every request with a long query — which is most of them.

Rerank is billed per token at the embedding rate; GET /v1/models lists it.


POST /v1/audio/speech — Pica

{"model":"pica-nano","input":"नमस्ते, आज मौसम बहुत अच्छा है।","voice":"david","mode":"natural"}

Returns audio/wav. The charge rides in headers, because the body is audio and a customer must be able to read what a synthesis cost without parsing a WAV:

X-Model: pica-nano    X-Voice: david        X-Mode: natural       X-Lane: standard
X-Cost-Paise: 1.1958  X-Cost-Currency: INR_paise
# and, when a request was moved: X-Overflow: pica-small (pica-nano was at capacity) or
# X-Fallback: pica-small (pica-large does not speak that language)
X-Duration-S: 2.05    X-Sample-Rate: 24000

Three tiers, billed per hour of audio

model What it is Price
pica-nano The eleven built-in voices, your own voices (mcv_…), the standard and expressive lanes, emotions, first_clause and chunk ₹21 per hour of audio
pica-small Its own Hindi and English voices ₹32 per hour of audio
pica-large The most natural of the three; every voice speaks Hindi and English ₹60 per hour of audio

⚠️ Billed per second of audio produced (since 0.45.0; before that Pica was one tier at ₹32 per 10,000 characters). The charge is the playing time of the audio Pica delivered to you, measured from the audio itself — ₹21 an hour is 0.5833 paise a second. A request that fails before any audio is not charged; a stream you hang up on is charged for the audio made up to that point. At most 5,000 characters per request.

  • When pica-nano is at capacity, a pica-nano request may be spoken by pica-small — only a built-in Hindi voice with no emotion; it is billed as pica-nano, and the answer says so with X-Overflow: pica-small and the X-Voice that spoke. Your own voices and emotions are never moved; they wait for pica-nano.
  • A language pica-large does not speak is spoken by pica-small (0.62.3). It speaks Hindi, English, Marathi, Tamil, Telugu, Gujarati, Bengali and Kannada; a text in Malayalam, Punjabi, Odia or Assamese is served by pica-small with the voice of the same name where it has one, and the tier's default otherwise. It is billed as pica-small, and the answer says so: X-Model: pica-small and X-Fallback: pica-small. Nothing is refused and nothing is read in the wrong language.
  • The language comes from the script you send (0.62.3). Before that, pica-large read every text that was not Devanagari as English, so Tamil, Telugu, Bengali, Gujarati and Kannada were read with an English tongue.
  • pica-mini is accepted as pica-small — the tier's name for its first day (0.45.0); the answer's X-Model says pica-small.
  • No model is pica-nano. ⚠️ One of the eleven voices — or one of yours — is always spoken by pica-nano and billed as pica-nano, whatever model you name: {"model":"pica-small","voice":"david"} keeps working as it did before the tiers existed, and X-Model: pica-nano says so.
  • pica-small and pica-large speak with their own voices — GET /v1/audio/voices?model=pica-small lists them. With no voice, pica-small uses ira for Indian-script text and zara otherwise, pica-large uses saanvi.
  • The pica-small and pica-large voices were renamed on 0.63.0 (23 September 2026): every pica-small and pica-large voice has a new id and name, and GET /v1/audio/voices lists only those. The old ids keep working — a request that names one is spoken by the same voice and answered with the new id in X-Voice. Nothing else about the request changes.
  • On pica-small and pica-large, lane, emotion, chunk and first_clause are 400 option_not_supported, not ignored. mode is accepted and changes nothing. stream and response_format: "mulaw_8k" work as below.
  • 429 tier_busy — the tier is at capacity right now; retry shortly or use pica-nano. 503 tier_not_deployed — that tier is not configured on this deployment.
Field
model pica-nano (default), pica-small or pica-large — see the table above
voice a voice of that tier, a built-in or one of yours (mcv_…) — GET /v1/audio/voices lists them all. Omitted on pica-nano, you get david for Devanagari text and robert otherwise
mode normal or natural (default). Delivery style, not the Osprey speed/intelligence ladder. On pica-nano the speech is paced after it is made: 1.30× for Hindi voices, 1.05× for English voices (since 0.61.4 — English at 1.30× ran about 270 words a minute)
lane standard (default, both languages) or expressive (Hindi, and the only one that takes an emotion)
emotion the expressive lane only — neutral · happy · sad · angry · disgust · fear · surprise
stream true streams length-prefixed WAV frames, one per unit (a sentence in natural mode) — see below
first_clause true speaks the first sentence's first clause on its own — a few words before the first frame instead of a whole sentence. Off by default; see First audio below
chunk "clause" speaks every sentence clause by clause. "stream" (stream only, standard lane only) sends each sentence as several frames while it is still being made — see Inside a sentence below. Off by default
response_format "mulaw_8k" on a stream delivers 8 kHz G.711 μ-law frames for a phone line — see below. Any other value is accepted and ignored: the answer is 16-bit PCM WAV at the voice's own rate

Streaming speech

"stream": true sends each sentence the moment it is synthesised, so the first one can start playing while the rest are still being made. The body is a sequence of frames, each flushed on its own:

┌──────────────────────────────┬────────────────────────────────────────────┐
│ 8 bytes, unsigned big-endian │ that many bytes: one complete WAV file      │
│ length N                     │ (RIFF header, PCM 16-bit mono)             │
└──────────────────────────────┴────────────────────────────────────────────┘
  … repeated, then the connection's chunked body ends
  • One frame per unit. In natural mode a unit is a sentence (a sentence ends at . ? ! । or a newline; a run such as ... or ?! ends it once, at its last mark). In normal mode the whole answer is one frame, all at once, when the whole text is done. With first_clause or chunk a unit is a clause (below), in either mode. X-Units says how many frames a whole stream has, before the first one arrives.
  • Every frame is a whole, playable WAV. Each carries its own header, so a player can start on frame one without waiting for a total length. Each ends with the pause that follows its sentence already in the audio — play the frames back to back, with no gap of your own.
  • The sample rate is the same in every frame and is sent up front in X-Sample-Rate: 24000 on the standard lane, 44100 on expressive, 24000 on pica-small and pica-large. The headers are the same as for a whole answer — X-Model, X-Voice, X-Mode, X-Lane — except X-Cost-Paise, X-Cost-Currency, X-Duration-S and X-Seed, which are not known until the end. The cost of a stream is on its usage row (GET /dashboard usage). Content-Type stays audio/wav; the framing is what this section describes.
  • There is no Content-Length and no end-of-stream frame: the stream is finished when the body ends. A body that ends inside a frame was cut off.
  • The same audio costs the same with or without stream: per second of audio. Hang up part-way and you pay for the audio made so far. Hanging up stops the synthesis as soon as your connection closes.
  • If synthesis fails before the first frame, the answer is a JSON error (502 speech_failed, or 503 through the public host), not charged. After the first frame the status is already 200; a stream that fails later simply has fewer frames than sentences, or ends inside a frame.
  • The stream is sent with Cache-Control: no-cache and X-Accel-Buffering: no, and its connection closes when it ends (Connection: close) — open a new connection for the next request.
  • The frames of a stream are not byte-identical to the one WAV the same request returns without stream: a whole answer is levelled and paced as one piece, a stream sentence by sentence.
  • With "chunk": "stream" a unit is several frames, and the stream says so with X-Chunk: stream — see Inside a sentence below.

First audio: first_clause and chunk

A whole first sentence has to be synthesised before the first frame can leave. Measured on the public host (2026-09-14, N 60): the gateway's time is about 530 ms + 15 ms per character, so a 50-character Hindi sentence is ~1.3 s to the first byte. "first_clause": true has Pica speak the first sentence's first clause — the words up to the first , ; : — or । that is followed by a space, plus following clauses while the head stays within about six words — as a frame of its own, then the rest of the sentence, then every later sentence exactly as it would have been. Measured with the same clause cut made client-side (N 56 pairs): the first audio arrived 437 ms earlier at p50, 1,007 ms at p90.

{"input":"आज शाम चार बजे, राजेश के साथ आपकी कॉल तय है। पॉइंट्स अभी तैयार हो जाएँगे।","voice":"david","stream":true,"first_clause":true}

→ X-Units: 3 and three frames: आज शाम चार बजे, · राजेश के साथ आपकी कॉल तय है। · पॉइंट्स अभी तैयार हो जाएँगे।.

  • A first sentence with no clause mark is not cut: you get the plan you would have had, and X-Units says so. A colon inside a time (10:30) or a comma inside a number (1,000, 1,25,000) is not a clause mark.
  • A clause is never spoken alone when it is under three words or ten characters — it is joined to the next one (नमस्ते, आज शाम… is not cut at all). Measured on these voices: a greeting generated on its own was swallowed or garbled, while the same words inside a longer take were perfect.
  • A very short sentence — fewer than 3 words or under 10 characters, such as नमस्ते! or Okay. — is joined to the next one (or to the previous one when it is last), for the same reason. This happens on every request, with or without these fields.
  • A full stop inside a number or after an abbreviation is not a sentence end, with or without these fields: 10.5 लाख, डॉ. शर्मा, ए.के., Dr., Rs. 500, No. 7, e.g. stay inside their sentence. After a number the same letters close the sentence (took 300 ms., paid 500 Rs.), and a.m./p.m. close it when a capital letter follows. Main St. The… is still read as one sentence.
  • The clause frame ends in a short breath (about 150 ms after pacing) rather than a sentence pause.
  • Every later sentence is unchanged — same seed, same audio as without first_clause. Only the first sentence is a different take, spoken in two pieces; whether the seam is audible is something to judge by ear on your own voice, which is why this is opt-in.
  • "chunk": "clause" does the same for every sentence: every clause is its own frame. first_clause is then redundant.
  • Both fields work without stream too — the one WAV is then joined from the same units — but the point of them is the first frame of a stream.
  • The charge does not change: it is per second of audio.

Inside a sentence: chunk: "stream"

first_clause still waits for a whole clause to be synthesised. "chunk": "stream" does not wait for the unit at all: Pica decodes the sentence while its speech is still being generated and sends each piece as it is decoded, so the first frame is the first fraction of a second of the sentence.

{"input":"आपने पिछले हफ्ते टू बीएचके फ्लैट के बारे में पूछा था। पॉइंट्स अभी तैयार हो जाएँगे।","voice":"david","stream":true,"chunk":"stream"}
Content-Type: audio/wav     X-Sample-Rate: 24000     X-Units: 2     X-Chunk: stream
  • X-Units still counts units (sentences, or clauses with first_clause) — it does not count frames. How many frames a unit becomes is decided while it is decoded and is not known up front (about four for a 50-character Hindi sentence). X-Chunk: stream is how you know frames are not units.
  • Every frame is still a whole, playable WAV. Each carries a 16-byte RIFF chunk rkfc between fmt and data — little-endian u16 unit (0-based, in order), u16 seq (0-based within the unit), u8 flags (bit 0: the unit's last frame), 3 reserved bytes. Any WAV reader skips it; you need it only to know where a sentence ends.
  • Play the frames back to back, with no gap of your own: the cut between two frames of a sentence is not in the audio. Only a unit's last frame ends in its pause.
  • A whole stream is X-Units frames with the last bit. A stream that ends after a frame without it ended inside a sentence.
  • The words are the same take as without chunk: "stream" (the same speech tokens, measured on 60 Hindi and 60 English sentences), decoded in pieces while they are made. It does not sound identical to the whole-sentence frame: the duration matches within about 1.5 %, the joins between pieces were not measurable in the delivered audio and transcripts were as accurate, but the timbre differs slightly — judge it by ear on your voice.
  • The level of a sentence is decided before the sentence exists, where a whole-sentence frame is levelled exactly. Measured: within about 0.5 dB on average and 2 dB at most on a Hindi voice; on an English voice 1.3 dB quieter on average and up to about 4.7 dB. Consecutive sentences can therefore differ in level by a few dB, and a voice you created, which we have not calibrated, can be further off.
  • Same request, same bytes, whatever the load — as long as the service's streaming configuration is unchanged; when we retune it, where the frames are cut changes and so do the bytes.
  • A sentence too short to be worth cutting (roughly under 36 speech tokens, ~1.4 s of audio) arrives as one frame.
  • 400 chunk_needs_stream without "stream": true; 400 chunk_not_supported on the expressive lane, which cannot decode inside a sentence; 400 chunk_not_available when Pica cannot stream inside a sentence right now — send the same request without chunk and you get sentence frames.
  • ⚠️ It costs Pica more time per sentence (a sentence is decoded several times while it is made), so under load other requests wait a little longer for their turn.
  • The charge does not change: it is per second of audio.

For a phone line: response_format: "mulaw_8k"

Telephony wants 8 kHz μ-law. On a stream, "response_format": "mulaw_8k" delivers exactly that:

{"input":"आज शाम चार बजे, राजेश के साथ आपकी कॉल तय है।","voice":"david","stream":true,"first_clause":true,"response_format":"mulaw_8k"}
Content-Type: audio/basic     X-Sample-Rate: 8000     X-Encoding: mulaw     X-Units: 2
  • The framing is the same — an 8-byte big-endian length, then that many bytes — but the payload is raw G.711 μ-law, one byte per sample, 8,000 per second, no header. Feed a frame to the line 160 bytes (20 ms) at a time as it comes.
  • Each frame is Pica's 24 kHz (or 44.1 kHz on expressive) WAV frame resampled with a windowed-sinc low-pass (measured: a 1 kHz tone survives at ≥ 60 dB SNR, anything above the 4 kHz Nyquist is removed by ≥ 60 dB) and μ-law encoded (≈ 38 dB SNR, which is μ-law). Frames are resampled on their own, which is safe: every frame fades in and ends in its own pause.
  • With "chunk": "stream" the frames of one sentence are resampled as one signal — byte for byte what the whole sentence in one frame would give — so no cut between them is audible on the line. The last ~8 ms of a frame wait for the next one, and the sentence's last frame flushes them. The μ-law frames carry no tag (they have no header to carry it in); X-Units and X-Chunk: stream are sent as for WAV.
  • ⚠️ With "chunk": "stream", each unit ends with an EMPTY frame — an 8-byte length of 0 and nothing after it. It carries no audio (feeding it to the line plays nothing); it is how you tell where a sentence ends. A whole stream has exactly X-Units empty frames and ends with one; a stream with fewer ended inside a sentence.
  • Without stream, mulaw_8k is a 400 format_needs_stream: a phone line is fed as the frames come, and a caller who asked for μ-law must never be handed a WAV at HTTP 200.
  • On pica-small and pica-large the audio is made at 8 kHz directly, so nothing is resampled.
  • The charge is the same as for WAV: per second of audio.

⚠️ emotion on the standard lane is a 400, not an ignored field. A flat reading of something you asked to sound angry, billed as if it worked, is worse than a refusal.

⚠️ An English voice will not read Devanagari — 400 script_mismatch. Measured: it says nonsense, and nonsense billed as speech is worse than a refusal.

GET /v1/audio/voices

Everything this key may speak with: the eleven built-in voices and the voices this account has made (both pica-nano), and the voices of pica-small and pica-large. Every voice names its model; ?model=pica-large lists one tier. A pica-large voice lists "languages":["hi","en"] — its language follows your text. Also the two delivery modes, the two lanes and which of them has emotions (all three pica-nano only). It never returns a voice's reference clip or its hash: those name a file on our storage and are the thing an abuser would want.

{"object":"list","built_in_count":11,"own_count":1,
 "data":[{"id":"david","name":"David","language":"hi","gender":"male","built_in":true,"model":"pica-nano"},
         {"id":"neha","name":"Neha","language":"hi","built_in":true,"model":"pica-small"},
         {"id":"archana","name":"Archana","languages":["hi","en"],"gender":"female","built_in":true,"model":"pica-large"},
         {"id":"mcv_k4t2…","name":"Anil","language":"hi","kind":"generated","duration_s":14.2,
          "created_at":"2026-09-11T04:10:00Z","created_by_key":"…","built_in":false,"model":"pica-nano"}],
 "modes":["normal","natural"],"lanes":["standard","expressive"],"own_voice_lanes":["standard"]}

POST /v1/audio/voices — make a voice

Two ways in, and they are not the same thing.

You send What happens Use it for
{"description": "warm, unhurried, mid-forties"} as JSON a new voice is generated from your description. It copies nobody — it invents a speaker who has never existed a brand voice, a narrator, anything where you want a person who is not a person
a reference recording as multipart/form-data your clip becomes the voice, cloned a voice you own and want to keep using
# generate
curl https://api.minicrow.com/v1/audio/voices -H "Authorization: Bearer mc_..." \
  -H "Content-Type: application/json" \
  -d '{"name":"Anil","description":"warm, unhurried, mid-forties","language":"hi"}'

# clone
curl https://api.minicrow.com/v1/audio/voices -H "Authorization: Bearer mc_..." \
  -F file=@reference.wav -F name=Anil -F language=hi -F attest=true

Answers 201 with the voice. Its id is then just a voice on POST /v1/audio/speech.

Field
description JSON only. A sentence or two, at most 600 characters. Describe the voice, not the words
file multipart only. 10 to 20 seconds of clean speech works best; 3 to 30 is accepted. WAV, MP3, FLAC, OGG, AIFF or CAF — m4a, aac and webm are not read
attest multipart only, and required: attest=true
name what you want to call it. Yours; we never show it to anyone else
language hi (default) or en. It chooses the pronunciation lane. hi reads both scripts; en refuses Devanagari, exactly as a built-in English voice does

What cloning a voice means here, stated plainly

⚠️ attest=true means you are saying: I hold the rights to this recording and have the speaker's permission to clone their voice. Sending it is a claim you are making, on the record, with a timestamp and the address it came from.

⚠️ MiniCrow does not verify that claim, and cannot. Nobody can tell from a clip whether the person in it agreed. There is no consent check behind this endpoint, and calling the attestation a safeguard would be selling you one that does not exist. What is real is this:

  • every attempt — accepted and refused — is logged against the API key that made it;
  • the reference recording is kept, so a complaint about a voice can be answered with the audio it was made from;
  • creation is rate-limited per key (10 an hour, 30 a day by default), and an account holds at most 100 voices at a time.

Whose voice it is

⚠️ A voice belongs to the account that made it and is never visible to any other customer. Another customer cannot list it, speak with it, or delete it — and an id that is not yours answers 404 unknown_voice, worded identically to a typo, because a 403 would confirm the id exists. The id itself is 128 random bits, so it cannot be guessed or walked.

⚠️ Every key on your account can use every voice on your account. The key that created a voice is recorded and is what the rate limit and any complaint are measured against — but rotating a key does not lose you your voices.

⚠️ A voice you made speaks on the standard lane only. The expressive lane has its own fixed cast, so a custom voice with lane: expressive is a 400 lane_not_supported rather than a stranger's voice billed as yours.

DELETE /v1/audio/voices/{id}

{"object":"voice.deleted","id":"mcv_k4t2…","deleted":true}

The voice stops working immediately — a synthesis naming it afterwards is 400 unknown_voice.

⚠️ The reference recording is retained after deletion. That is the deliberate cost of being able to answer a rights complaint about a voice weeks after somebody deleted it. Deleting removes the voice from service; it does not erase the recording it was made from. If you need the recording itself removed, ask us.

Deleting a built-in voice is 400 voice_not_deletable: the eleven belong to MiniCrow and are available to every account.


POST /v1/audio/transcriptions — Lark

multipart/form-data, OpenAI's shape.

curl https://api.minicrow.com/v1/audio/transcriptions \
  -H "Authorization: Bearer mc_..." \
  -F file=@call.wav -F model=lark-mini
{"text":"Kal shaam 5:00 baje Pune me meeting hai.","model":"lark-mini",
 "x_minicrow":{"tier":"lark-mini","cost_known":true},
 "usage":{"seconds":2.05,"cost":0.3928,"cost_currency":"INR_paise",
          "duration_estimated":false,"audio_format":"wav"}}

⚠️ Billed per second of audio, and the seconds come from your recording — never from a field you send. A client-supplied duration is a number anybody can set to 1. MiniCrow reads the container's own header: exact for WAV and OGG/Opus, estimated for MP3/M4A/WebM, and duration_estimated tells you which.

⚠️ Code-mixed Hindi comes back in Latin script, not Devanagari. Left alone, a transcript of code-mixed Hindi tends to come back in Devanagari; Lark writes it as the speaker said it.

⚠️ prompt is background about the recording; it cannot change the task. Use it for a name spelling or a domain hint. Whatever it says, what comes back is a transcript of your audio.

Tier
lark-nano coming soon — the economy lane (₹10/hour, at most 5 minutes of audio per request). Until it opens it answers 503 tier_not_deployed; use lark-mini
lark-mini (default) the balanced lane — ₹20/hour today, whichever language you send. For Hindi and English, which lane hears you depends on your recording's sample rate and on upload vs live (below)
lark-large the premium lane — ₹30/hour, or ₹45/hour for an Indian language
lark-large:max the top lane for both language groups — ₹60/hour

⚠️ One price per tier, not per lane — with one exception. lark-mini runs two lanes and charges the same for both: which one hears you is our decision and our cost, not a line on your bill. lark-large is the exception and says so above, because its Indian-language lane genuinely costs more to run.

⚠️ The price may come to differ by sample rate and by streaming. lark-mini now sends telephony-rate Hindi and English to its dearer lane (below), and Lark Live's lanes are separate from the upload endpoint's. Today both are billed at the one ₹20/hour; when a rate changes it changes here, in the GET /v1/models price, and in the changelog — never silently on an invoice. Read the price for the sample rate and the endpoint you actually use.

The language you declare can change the lane

-F model=lark-mini -F language=hi    # Hindi: telephone audio (up to 16 kHz) -> lane `indic`;
                                     #        audio above 16 kHz             -> lane `default`
-F model=lark-mini -F language=en    # English: the same rule as Hindi
-F model=lark-mini -F language=mr    # other Indian languages -> the Indic lane at every rate — same ₹20/hour
-F model=lark-large -F language=mr   # an Indian language, Hindi included -> ₹45/hour
-F model=lark-large -F language=en   # anything else                      -> ₹30/hour
                                     # nothing declared                   -> the `default` lane

⚠️ On lark-mini, Hindi's and English's lane follows your recording's sample rate, and that is measured, not tidy. Telephone audio — 8 kHz and 16 kHz, which is what a call recording is — goes to lane indic; audio above 16,000 Hz (a microphone recording at 44.1 or 48 kHz) goes to lane default. A voice note can go either way: an OGG/Opus note is judged by the input rate its header declares, and many phone apps (WhatsApp among them) record voice notes at 16 kHz, so those take lane indic; check audio.sample_rate_hz in the response. On 100 Hindi clips with human-written references pushed through a telephone channel, the indic lane was a full point better in word error rate and the default one returned nothing at all on four of the hundred; on 100 English clips through the same channel the gap was far wider. At 48 kHz the two had measured level on Hindi. The line is 16,000 Hz: at or below it is telephony. It is a rule we can move without a release, so the numbers here are today's. lark-large has no such rule — Hindi is an Indian language there at every rate, and English takes its default lane.

⚠️ indic is the name of a lane, not a claim about your language. On lark-mini it is the lane that hears telephone audio best, and English telephone audio is sent to it for that reason. x_minicrow.lane: "indic" on an English call is expected.

⚠️ The rate is read from your file's own header, never decoded, and a rate we cannot read takes the telephony lane. WAV (fmt ), OGG/Opus (the encoder's input rate — Opus itself always runs at 48 kHz, so the header's input rate is what your recorder was fed) and MP3 (the frame header) are read. M4A and WebM are not opened that far: Hindi and English in those containers take lane indic — the language's base rule — and the response carries no audio.sample_rate_hz, so you can tell "not read" from "read and below the line". Send WAV or OGG if you want the rate to decide.

⚠️ Live is telephony by construction. The live socket takes 8 and 16 kHz only, so live Hindi always uses the Indic lane's final pass; the lane is hindi and its price is in the live section.

⚠️ The same file can therefore cost differently depending on how you send it. Today every one of these is ₹20/hour; the price is set per lane and per endpoint, so it may come to differ by sample rate and by streaming. Check GET /v1/models and this page for the endpoint and the rate you actually use.

⚠️ You declare it; we do not guess. Language cannot be detected before transcription, and a wrong guess sends Marathi to a lane that has never heard it. A request with no language takes the default lane — which is also the cheaper of the two to get wrong on lark-large.

The response says which one served and why, so you can tell without asking:

"x_minicrow": {"tier":"lark-mini","lane":"indic","lane_reason":"language",
               "audio":{"container":"wav","channels":1,"sample_rate_hz":8000}}
lane_reason
mode you named a mode (lark-large:max)
language the language you declared chose the lane
default no rule applied; the tier's default lane
sample_rate the language's rule has a second lane above a sample rate, and your recording was above it — today, Hindi and English on lark-mini above 16 kHz

audio.sample_rate_hz is the rate that was read; it is absent when the container did not say, and then the language's base lane served.

Indian languages: hi mr gu bn ta te kn ml pa or as ur ne sa kok mai sd ks doi mni sat brx bho raj (on lark-mini, hi — and en — by sample rate as above).

The alphabet is yours to choose

                     # nothing sent   -> the language as it is WRITTEN. The default.
-F script=native     # Devanagari for Hindi and Marathi, each language in its own script
-F script=latin      # romanised — "kal shaam ko meeting hai"
-F script=auto       # whatever Lark writes, with no alphabet asked for either way

devanagari, original and source are synonyms for native; roman and romanised for latin. Anything else is a 400 unknown_script rather than a silent fall back — a caller who typed devnagari and received romanised text could not tell that from an API that ignored them.

⚠️ native keeps your code-mixing, it does not translate it into Devanagari. An Indian business call is often half English: the speaker says meeting, link, online, payment inside a Marathi sentence. Those stay in English, in Latin letters, exactly as spoken — you get उद्या meeting आहे का?, never उद्या मीटिंग आहे का?. Numbers a speaker says as numbers come back as digits. Nothing is translated.

The response tells you which alphabet you asked for and what happened:

"x_minicrow": {"script": "native", "script_honoured": true, "script_repaired": false}

⚠️ A language written in the Latin alphabet is already in its own script. For en, fr, es and the other Latin-written languages, native means Latin: an English transcript is script_honoured: true and is never rewritten. With no language declared, an all-Latin transcript is not judged (script_honoured absent) — English and romanised Hindi look the same, and we will not guess. So a romanised Hindi transcript is only rewritten into Devanagari when you declare language=hi (or another Devanagari-written language).

⚠️ script_repaired: true means we had to fix it, and you were not charged for that. About one Indian- language recording in seven is first transcribed in the wrong alphabet — the words correct, the letters not — and retrying reproduces it exactly. Rather than hand you a writing system you did not ask for, the transcript is rewritten into the one you did. It is a script conversion and nothing else: no word is translated, corrected, added or removed. Both ways it can go wrong are repaired: English words written in Devanagari go back to English letters, and Hindi or Marathi written in English letters (make sure karna ki) goes into Devanagari (make sure करना कि). A repair that would change the words is refused; the transcript then comes back as it was, with script_honoured: false.

Who spoke — speaker labels

-F diarize=true
-F speakers=2      # optional: how many people are on the recording, if you know

Adds ₹3.50 per hour of audio on top of the tier's own rate, and returns the turns beside the transcript:

{
  "text": "Hello sir, mi Pooja bolte Samarth Sky project madhun. Ha bola, kay aahe sanga...",
  "segments": [
    {"start": 0.0,  "end": 4.2,  "speaker": 1},
    {"start": 4.5,  "end": 7.1,  "speaker": 2}
  ],
  "x_minicrow": {"speakers": 2}
}

⚠️ The transcript is unchanged. Labels do not get interleaved into text — that would break every caller already parsing it, and would throw away the timings, which are the half a CRM actually needs. Join them yourself on the times.

⚠️ Tell us speakers when you know it. A two-party phone call is two people; leaving Lark to work that out costs accuracy you could have had for free.

⚠️ If the speaker pass fails, you still get the transcript and you are not charged for the labels. The words were already correct and already earned; segments comes back empty and the ₹3.50 is not applied.

How good is it? Measured on constructed two-speaker telephone audio with exact ground truth: 100% of frames attributed to the right speaker, three recordings, and roughly 25× faster than real time. That is a ceiling, not a promise — those were two clearly different voices with little overlap. Real calls have similar voices, crosstalk and a worse line; published figures for speaker labelling of this kind land nearer 85–90%, and we have not yet measured it on a hand-labelled real call.

The single biggest thing you can do is record each party on their own channel. Then there is nothing to infer — the channel is the speaker — and it costs nothing.

Two channels, two speakers

An upload may hold one speaker or many, and we do not assume a two-channel file is two speakers. A two-channel WAV (16-bit PCM, μ-law or A-law) is transcribed channel by channel only when the file shows it: each channel carries speech of its own — energy above 1 kHz, a spectrum that keeps moving, loudness that rises and falls at the rate of syllables — and the two channels are not the same signal (correlation below 0.9 at every alignment within ±50 ms, so a copy of one channel written a few milliseconds late counts as the same signal). A silent channel, a steady tone or line hum on one side, or the same mix written to both channels fails, and the file is transcribed as one stream exactly as a mono file is.

⚠️ What the check cannot tell apart. It measures whether each channel sounds like speech, not whose speech it is, so: music on one channel (hold music with a beat) can pass at 8 and 16 kHz, and steady background hiss can pass too — that channel is then transcribed on its own and comes back empty or near-empty. And a panned mix — both voices on both channels at different levels, e.g. one voice at full level and the other at half — correlates below 0.9 and splits, so each channel's transcript repeats most of the conversation. If your recorder pans rather than separates, send a mono file. Compressed two-channel files (OGG, MP3, M4A, WebM) are not opened that far and always take the one-stream path.

When it splits, both channels go to the same lane, and the response labels them:

{
  "text": "channel_1: Hello sir, mi Pooja bolte...\nchannel_2: Ha bola, kay aahe...",
  "segments": [
    {"start": 0.0, "end": 64.2, "channel": 1, "speaker": "channel_1", "text": "Hello sir, mi Pooja bolte...",
     "script_honoured": true, "script_repaired": false},
    {"start": 0.0, "end": 64.2, "channel": 2, "speaker": "channel_2", "text": "Ha bola, kay aahe...",
     "script_honoured": true, "script_repaired": false}
  ],
  "x_minicrow": {"channel_split": {"split": true, "correlation": 0.12,
                 "rule": "each channel was transcribed on its own on the same lane; audio seconds are billed once"}}
}

⚠️ Channels count from 1: channel 1 is the left channel, channel 2 the right, in text, segments and channel_split's reasons alike.

⚠️ The alphabet is judged channel by channel. Each segment carries its own script_honoured and script_repaired; x_minicrow.script_honoured is true only when every channel that can be judged is, and script_repaired is true when any channel was repaired.

⚠️ Billed once, by the length of the recording. A 64-second two-channel call is 64 billed seconds, not 128.

⚠️ Each segment spans the whole call. A channel's words are not cut into timed turns; speaker is a channel label, a string, where the speaker pass (below) returns a number.

⚠️ diarize=true on a file that splits is skipped, not charged, and not refused. The channels already separated the speakers; x_minicrow.diarization says {"skipped": true, …}. On a file that does not split, diarize works exactly as described above.

When a two-channel WAV does not split, x_minicrow.channel_split says why: {"split": false, "correlation": 0.998, "reason": "both channels carry the same audio; transcribed as one stream"}. A mono file has no channel_split at all.

⚠️ Live is one speaker per stream. The live socket takes one channel; for a two-party call open one session per leg.

The transcript is of your recording, never one of Lark's own lines

On a recording with little or nothing in it, Lark can answer with one of its own reference lines instead of a transcript. MiniCrow checks every transcript against those lines; a transcript that is one — at least 80 % of the line's words in a row, making up at least 80 % of the transcript — is retried once, and if the retry does the same, the request fails with transcription_failed (HTTP 503 at our edge) and nothing is charged — the line never reaches you. A transcript that merely contains such a line's words, such as a real call that opens with the same line and carries on, is returned as normal. ⚠️ A recording whose entire content is one of those lines, word for word, is indistinguishable from a copy and is refused.

Declaring the language improves accuracy

0.35.3 (pending). Lark's default handling is tuned for Hindi/English code-mix and for Marathi. That helps Hindi and Marathi, and it hurts other languages: on other Indian languages Lark sometimes wrote the transcript in Devanagari, or in Marathi, or romanised it. So when you declare one of these twelve languages, Lark transcribes it as that language alone, word for word and never translated:

language what you get (with script=native or no script)
ta te bn gu kn ml the language in its own script (Tamil in Tamil script, and so on); English words the speaker says stay in English
en es de fr it ar the language as it was spoken, not translated

Measured on 100 FLEURS clips per language through an 8 kHz telephone line, on the indic lane: character error fell by 16.4 points on Kannada, 8.8 on Gujarati, 5.8 on Bengali, 5.6 on Tamil, 4.7 on Malayalam and 3.5 on Telugu, and word error fell by 1.8 on Arabic. On all of these, no clip came back in the wrong alphabet (38 of 600 Indian-language clips did before). English, Spanish, German, French and Italian scored the same either way.

  • hi, mr, no language, and every other language keep the default handling, unchanged. For Hindi and Marathi, the per-language handling wrote English words in Devanagari (फ्लॅट), and your code-mixed calls need them in Latin. That has to be checked on real calls before those two change.
  • script=latin on Tamil, Telugu, Bengali, Gujarati, Kannada, Malayalam or Arabic asks for the Latin alphabet (romanised) instead of the language's own script. script=auto names no script at all. On English, Spanish, German, French and Italian, latin is the same as native. Neither variant has been measured.
  • domain, vocabulary, abbreviations and prompt work the same way with these languages. For a language whose script is named, the script you asked for always has the last word.
  • The copy check above and the alphabet repair are the same for every language.

Language packs

0.37.0 (pending). Switched on lane by lane; until your lane is switched, everything above applies unchanged.

A language pack tunes Lark to one declared language alone — its script, how English words are written in it, its numerals and punctuation — and to nothing about any business. Packs exist for 14 languages:

language how English words the speaker says are written (script=native or no script)
mr hi bn ml in the language's own script, as a newspaper in that language prints them (बस, ऑफिस)
ta te gu kn in English letters; a borrowed word that is part of everyday speech stays in the language's script
en es de fr it ar the language's own spelling and punctuation, numbers as digits; Arabic without added vowel marks

⚠️ For mr, hi, bn and ml this changes what your transcript looks like. Without a pack an English word such as "flat" comes back in Latin letters inside a Devanagari sentence; with the pack it comes back in Devanagari. If you store or search code-mixed text with English words in Latin letters, ask for script=latin (which keeps today's behaviour) or tell us before your lane is switched.

  • Only with script=native or no script. script=latin on an Indian language, script=auto, no language, and every language without a pack keep the behaviour above, unchanged. On en es de fr it, latin is the same as native.
  • Measured on 50 held-out read-speech clips per language through an 8 kHz telephone line: without the packs, Lark trailed a leading Indian speech-to-text service by 1.75 words per hundred pooled over nine languages; with the packs it was level with it in every one of those languages. ⚠️ That measurement also gave each clip a draft from Lark Live, which this endpoint does not do. Without the draft, the pack set that ta te gu kn en es de fr it ar use was ahead of no pack by 0.43 words per hundred, pooled over 700 clips in all 14 languages (not held out, and only just outside the margin); the set mr hi bn ml use has not been measured without the draft.
  • The copy check compares your transcript with the pack's own reference lines, and only those. The alphabet repair is not applied: script_honoured is true when the transcript carries letters of the language's own script (for en es de fr it, when it carries no other script).
  • domain, vocabulary, abbreviations and prompt work with a pack as they do without one, and never override it.
  • On Lark-V, the speech half uses the pack when the Lark-V tier is switched, independently of the audio tiers.

Tell us about your own recording

Lark's own handling is the same for every customer: it knows your language and script and nothing about your business. What your recordings are about comes only from you, in these optional fields:

-F domain="<who is talking to whom, about what — one line>"
-F vocabulary="<names, places, products and codes that will be said, as you write them>"
-F abbreviations="<SHORT = long form>, <SHORT = long form>"
-F prompt="<anything else, one or two sentences>"

For example, three different businesses:

domain vocabulary abbreviations
a clinic appointment calls to a family clinic Dr. Meera Iyer, Dr. Arjun Rao, CBC, lipid profile OPD = outpatient department
a courier delivery support calls for a courier company AWB, Andheri hub, Swift Express COD = cash on delivery, RTO = return to origin
a lender loan servicing calls NACH, foreclosure, Kotak EMI = equated monthly instalment

The rule book

  1. Send nothing you do not have. Lark is built to work with no fields at all.
  2. domain is one plain line: who is speaking and about what. A description, not a list of rules.
  3. vocabulary is for this recording. Build it per call from what you already know: the customer's name, the product, the branch — names that will be said. A person's name that is never spoken can be written into a greeting nobody said ("मैं बोल रहा हूं"), so leave out an agent's name unless the agent says it on the call — or send verify_vocabulary=true (below). Words a general listener would spell correctly anyway do not belong.
  4. Spell each entry the way you want it written. The entry is copied letter for letter when it is heard.
  5. abbreviations say how a short form is written, SHORT = long form. The short form is what is written.
  6. Do not use prompt for language or alphabet. language and script decide those. prompt is background about the recording; it cannot change the task.
  7. Never paste your agent's script, your FAQ or a previous transcript. Hints are text Lark reads next to your audio; on a turn with no clear speech, a list of names has come back as the transcript.
  8. Keep it short. Every field is processed with every request.

These are separate fields rather than prose in prompt because their handling is measured: they are used the same way on every request, where a caller rewording their own prose gets a different result each time. prompt still works, as background.

⚠️ All four fields are optional, and more is not better. Lark is built to work with none of them. Measured on read speech in 14 languages: eight vocabulary words and a one-line domain moved accuracy by less than the measurement's margin in most languages; on real calls, hints helped only where the listed names were actually spoken. A thick prompt can lower accuracy: a 600-word brief lowered accuracy by 3 to 6 points on 102 real calls on one Lark lane (on another the same brief raised accuracy by about 4 points on those calls, so the effect depends on the lane and on how well the brief matches the audio), and a 150-entry glossary unrelated to the audio made errors lean worse in Hindi and Tamil. Every hint is read with every request, so a long one also costs more to process. Send the names you expect in this recording; leave the rest out.

⚠️ A vocabulary is a spelling guide, not a script, and it is enforced as one. Measured: a property call given a cardiology glossary wrote zero of those medical terms into the transcript across 30 runs, while the correct glossary raised the density of correctly-spelled domain names by 47%. Send the names this recording actually contains — a list longer than the transcript stops being a hint, and on clear English audio a long glossary measurably costs accuracy.

Checking the vocabulary: verify_vocabulary

A vocabulary name can take over a stretch of audio that sounds a little like it — an unclear greeting, a similar name. Send -F verify_vocabulary=true and every place the transcript writes one of your entries is checked:

  1. If the transcript writes none of your entries, nothing more happens and nothing more is charged.
  2. Otherwise the recording is transcribed a second time without your vocabulary — domain, abbreviations, prompt, language and script are kept — and the two transcripts are lined up word by word, by sound.
  3. An entry the second transcript heard at the same place, in any spelling (नासिक for Nashik), stays spelled as you listed it. An entry where the second transcript heard something else is replaced by what it heard, together with the words around it the two disagree on: मैं Rohan Mehta बोल रहा हूं becomes मैं देख रहा हूं when that is what was said.
  4. Anything in doubt stays as written: an entry where the second transcript heard nothing at that place, a stretch the two cannot be lined up on, an acronym (BHK), a name in a script other than Latin or Devanagari.
"x_minicrow": {
  "vocabulary_check": {"ran": true, "confirmed": ["Nashik"], "removed": ["Rohan Mehta"], "unverified": []}
},
"usage": {"seconds": 47.1, "cost": 52.22, "vocabulary_check_cost": 26.11, "cost_currency": "INR_paise"}

Price: the second transcription is billed as one more transcription of the recording — its seconds, once, at the same tier's rate — only when it ran and answered. usage.vocabulary_check_cost is that part and is included in usage.cost. On a two-channel call only the channels that wrote an entry are transcribed again, and the recording's seconds are billed once more. When the check does not run, vocabulary_check is {"ran": false, "reason": …} (no vocabulary sent; no entry written; the second transcription failed; a lane that does not use a vocabulary, such as lark-large:max) and nothing extra is charged. On a split call each channel's segment carries its own vocabulary_check.

Measured: on 50 real telephone calls, each with the names spoken on it as the vocabulary, the check kept 129 of 132 spoken names written, left 3 as written, and took out none; the one name written over other words in the test set was replaced by what was said. It does not catch a listed name that sounds like the one said — Mehta written for a spoken मेहरा — because at that closeness it cannot tell a wrong name from a spelling. Off unless you ask for it.

⚠️ A recording with no speech is not transcribed, and costs nothing. A WAV that certainly carries no speech — silence, a steady tone or hum, flat line noise, a busy or ringing tone — is not sent to a model: it comes back with "text": "", x_minicrow.no_speech: true and usage.cost: 0. It is judged on the audio itself, for WAV files of 3 seconds or more; anything in doubt is transcribed as usual. On a two-channel call transcribed channel by channel, each channel is judged on its own: a channel that certainly holds no speech comes back as an empty segment with "no_speech": true and is not sent to a model, and when no channel holds speech the whole call is answered as above. Any other two-channel file is answered this way only when every channel is without speech. A tone that switches on and off — as some lines do while the other side talks — can still pass as speech and be transcribed: measured on 634 recordings, no rule we tried separated such a tone from real speech without also dropping real speech. x_minicrow.channel_split only says whether each channel looks like one speaker; it is not a no-speech verdict.

⚠️ A transcript that repeats itself is never returned. If the model falls into a loop — one sentence written over and over — the recording is transcribed once more without your domain, vocabulary, abbreviations and prompt, and that transcript is returned if it does not loop; otherwise the repeated words are kept once. Either way x_minicrow.repetition says so ("retried_without_hints" or "collapsed"); it is absent on every other response. It only fires on a run no speaker makes — one unit of words repeated back to back at least four times and at least twenty words long — and the recording is billed once.

Oversize is refused, never truncated: domain 120 characters, vocabulary 2,000 characters or 150 entries, abbreviations 1,200 characters or 100 entries, prompt 2,000 characters.

On every other tier language is accepted and does not move the lane, because there is only one lane to move to. The lane in the response always says which one served.

What this endpoint refuses

At most 25 MB per request, matching OpenAI's limit. Beyond that:

Status Code When
400 file_required no file part in the form
400 unsupported_audio the container is not one MiniCrow can measure. Send WAV, OGG/Opus, MP3, M4A or WebM — the charge is per second and the seconds come from the file's own header, so a format we cannot read is one we cannot bill. FLAC and raw PCM land here
400 empty_audio the file is a valid recording of zero seconds. It is refused rather than transcribed: a model handed silence will confidently write words nobody said, and a free transcription is indistinguishable from a working one until somebody reads an invoice
400 file_unreadable the part is empty or could not be read
400 unknown_model / unknown_mode no such tier, or a mode that tier does not offer
413 file_too_large over 25 MB; split the recording
413 audio_too_long lark-nano only — at most 5 minutes of audio in one request. 25 MB is a size limit, not a duration one: 25 MB of Opus is nearly an hour. Split the recording, or send it to lark-mini, which has no such limit
503 tier_not_deployed the tier, or the mode, is not serving on this deployment. The message names what to use instead

GET /v1/audio/transcriptions/live — Lark Live

Preview — enabled per account. Lark Live is open to named accounts while it is measured on real traffic. On any other account every connection is refused 503 tier_not_deployed. Ask us to enable yours.

A WebSocket. You stream a phone call's audio as it happens; MiniCrow finds where each speaker's turn ends and answers every turn twice — a fast draft the moment the turn ends, then a checked final once the turn has been listened to again. On intl languages you also get words while the turn is still being spoken.

wss://api.minicrow.com/v1/audio/transcriptions/live?model=lark-mini&language=mr&script=latin&encoding=mulaw&sample_rate=8000
Authorization: Bearer mc_...

⚠️ The key goes in the Authorization header, and nowhere else. Not in the URL — a query string is written into proxy and server logs — and not in a first message. A browser's WebSocket cannot set a header, so connect from your server, never from a web page: a key in a page is a key anybody can read.

Python

# pip install "websockets>=14"
import asyncio, json, os, wave
from urllib.parse import urlencode

import websockets

params = urlencode({
    "model": "lark-mini", "language": "mr", "script": "latin",
    "encoding": "linear16", "sample_rate": 16000,
    "domain": "delivery support calls for a courier company",
    "vocabulary": "AWB, Andheri hub, Swift Express",
})
URL = f"wss://api.minicrow.com/v1/audio/transcriptions/live?{params}"
HEADERS = {"Authorization": f"Bearer {os.environ['MINICROW_API_KEY']}"}


async def send_audio(ws, path):
    with wave.open(path, "rb") as w:                 # 16 kHz, mono, 16-bit PCM
        assert (w.getframerate(), w.getnchannels(), w.getsampwidth()) == (16000, 1, 2)
        while chunk := w.readframes(1600):            # 100 ms of audio
            await ws.send(chunk)                      # bytes are sent as a binary frame
            await asyncio.sleep(0.1)                  # real-time pace, as a live call arrives
    await ws.send(json.dumps({"type": "end"}))


async def main():
    try:
        async with websockets.connect(URL, additional_headers=HEADERS) as ws:
            sender = asyncio.create_task(send_audio(ws, "call-16k.wav"))
            async for message in ws:                  # ends when the server closes with 1000
                event = json.loads(message)
                if event["type"] == "transcript.partial":
                    print(f"  ({event['stage']}) {event['text']}")
                elif event["type"] == "transcript.final":
                    print(f"[{event['start']:.2f}-{event['end']:.2f}] {event['text']}")
                elif event["type"] == "session.ended":
                    print(f"billed {event['billed_seconds']} s, {event['cost']} paise")
                elif event["type"] == "error":
                    print("error:", event["code"], event["message"])
            await sender
    except websockets.exceptions.InvalidStatus as refused:   # refused before the socket opened
        print(refused.response.status_code, refused.response.body.decode())


asyncio.run(main())

JavaScript (Node)

// npm install ws — save as live.mjs, run with node live.mjs
import fs from "node:fs";
import WebSocket from "ws";

const params = new URLSearchParams({
  model: "lark-mini", language: "hi", script: "latin", encoding: "mulaw", sample_rate: "8000",
});
const ws = new WebSocket(`wss://api.minicrow.com/v1/audio/transcriptions/live?${params}`, {
  headers: { Authorization: `Bearer ${process.env.MINICROW_API_KEY}` },
});

// Refused before the socket opened: the body is the usual JSON error.
ws.on("unexpected-response", (req, res) => {
  let body = "";
  res.on("data", (chunk) => (body += chunk));
  res.on("end", () => console.error(res.statusCode, body));
});

ws.on("open", async () => {
  const audio = fs.readFileSync("call.ulaw");          // raw 8 kHz μ-law, no header
  for (let i = 0; i < audio.length; i += 800) {         // 800 bytes = 100 ms
    ws.send(audio.subarray(i, i + 800));                // a Buffer is sent as a binary frame
    await new Promise((resolve) => setTimeout(resolve, 100));
  }
  ws.send(JSON.stringify({ type: "end" }));
});

ws.on("message", (data) => {
  const event = JSON.parse(data.toString());
  if (event.type === "transcript.final") console.log(`[turn ${event.turn}] ${event.text}`);
  if (event.type === "session.ended") console.log(`billed ${event.billed_seconds} s, ${event.cost} paise`);
  if (event.type === "error") console.error("error:", event.code, event.message);
});

ws.on("close", (code, reason) => console.log("closed", code, reason.toString()));

Query

Everything is decided before the socket opens. A wrong parameter is an ordinary JSON error on the upgrade request, not a socket that opens and then closes.

Parameter Values
model lark-mini (default) the only tier with a live lane
language required. mr bn te ta gu kn or ml pa as · hi · en es fr de pt it ru ar ja ko a region is accepted and ignored — mr-IN is mr. Any other language is refused, never guessed
script latin or native — required for every Indian language, Hindi included the same synonyms as the batch endpoint (roman, devanagari, …). auto is refused on the live lane. Ignored for en es fr de pt it ru ar ja ko, which are written in their usual script
encoding linear16 (default) · mulaw · alaw linear16 is signed 16-bit little-endian mono PCM. μ-law and A-law are G.711, as a telephone line carries them
sample_rate 8000 · 16000 mulaw and alaw are 8000 only. Omitted, it is 16000 for linear16 and 8000 for mulaw/alaw
endpointing_ms 300–1000, default 400 how much silence ends a turn. Lower answers sooner and splits more sentences in two
vad server (default) · client client: you decide where a turn ends by sending {"type":"turn_end"}
domain at most 120 characters what the calls are about, one plain line — delivery support calls for a courier company
vocabulary at most 40 terms, 2,000 characters, comma-separated names as they are written. Only written into a transcript when they are said. 0.41.0: also boosted inside the draft on hi, mr and en — the languages the boost is measured on (see Hearing vocabulary under Osprey Live: names not words; for hi/mr add the Devanagari spelling as its own term). 0.43.0: hi and mr drafts are locked to Devanagari, vocabulary or not — English words in English letters, digits and punctuation stay
context repeatable, at most 3 values of 500 characters each the last finals of an earlier session, when you reconnect mid-call

The response says which lane serves you: indic for mr bn te ta gu kn or ml pa as, hindi for hi, intl for the rest. It is decided by language alone — as on the batch endpoint, you declare it and we do not guess.

⚠️ domain, vocabulary and context are optional. They are used on every turn, so a long vocabulary costs more on every turn and can add latency to every final; eight words and a one-line domain measured within 40 ms. Send names that will actually be said. A list of names can also be read back as a final on a turn with no clear speech. The rule book under Tell us about your own recording applies here too; on a live call, build the vocabulary per call (the customer's name, the product, the branch) rather than one list for every call.

0.37.0 (pending), language packs on live. When your live lane is switched, a session with script=native in hi ta bn gu ml (English words written in the language's own script) or mr te kn (English words in English letters) — and every session in en es de fr it — gets the same kind of domain-free language pack as the batch endpoint (see Language packs above). script=latin, and every other language, Arabic included, keep today's live behaviour. A session keeps the pack it started with, or none, to its end. The earlier finals it carries are the last three that were not empty.

⚠️ Speaker labels are not offered live yet. diarize or speakers on this endpoint is 400 diarize_not_available_live. Record each party on its own channel and open one session per channel (each is billed) — then the session is the speaker.

What you send

Message
binary frame audio in the encoding and sample_rate you declared, 20 ms to 1 s per frame. A frame over 64 KiB closes the session with 1009
{"type":"turn_end"} vad=client only: the turn in progress ends here
{"type":"end"} no more audio. Every open turn is finished, then session.ended, then close 1000
{"type":"ping"} answered with {"type":"pong"}

Any other text message is answered with a non-fatal error bad_message and the session carries on.

You may send a recording faster than real time. While 16 turns are still waiting for their finals we stop reading your socket until they are sent, so your writes slow down instead of piling up. One session carries at most 3 hours of audio, however fast it arrives.

⚠️ Send end, do not just hang up. end waits for the last turn's final. A dropped socket is still billed for every second it delivered, and the finals still in flight are lost.

What you receive

One JSON object per message, in this order for every turn: speech.started → transcript.partial → transcript.final.

{"type":"session.started","session_id":"6f1c…","model":"lark-mini","lane":"indic","language":"mr",
 "script":"latin","encoding":"mulaw","sample_rate":8000,"endpointing_ms":400,"vad":"server"}
{"type":"speech.started","turn":0,"start":0.42}
{"type":"transcript.partial","turn":0,"stage":"draft","text":"हो सर फ्लॅट चा बुकिंग अमाउंट पन्नास हजार आहे",
 "start":0.42,"end":3.18,"confidence":null}
{"type":"transcript.final","turn":0,"text":"Ho sir, flat cha booking amount 50 hazar aahe.",
 "start":0.42,"end":3.18,"script":"latin","lane":"indic","repair_ms":1180,"fallback":null,"script_honoured":true}
{"type":"session.ended","session_id":"6f1c…","audio_seconds":124.37,"turns":31,"billed_seconds":125,
 "cost":69.4445,"cost_currency":"INR_paise","cost_known":true}
Event Fields
session.started session_id, model, lane, language, script, encoding, sample_rate, endpointing_ms, vad
speech.started turn (from 0), start — seconds into the audio you have sent, not wall-clock time
transcript.partial turn, text, start, stage. stage: "streaming" — intl only, repeated while the turn is spoken, each one the whole turn so far. stage: "draft" — every lane, exactly once, when the turn ends; it adds end and confidence (null when not measured)
transcript.final turn, text, start, end, script, lane, repair_ms, fallback, script_honoured
session.ended session_id, audio_seconds, turns, billed_seconds, cost, cost_currency: "INR_paise", cost_known
error code, message, fatal — after a fatal error the session ends and the socket closes
pong —

⚠️ Every speech.started gets exactly one transcript.final, and finals arrive in turn order. A later turn can be ready first; it is held until the turn before it is sent, so you never have to reorder. A turn that held no words — a cough, a line click — still gets its final, with text: "" and fallback: null.

⚠️ The draft is fast and rough; the final is the transcript. Use the draft to react while the caller is still on the line, and store the final. fallback: "draft" means the check could not complete in time and the final is the draft — you still get a final for the turn, never silence. fallback: "unavailable" means the session ended on our side before the turn could be transcribed at all: text is "", but someone may have spoken.

⚠️ script applies to the final, not the draft. An Indian-language draft is written as it was first heard — usually in the language's own alphabet, as in the example above — even when you asked for latin. Only the final is converted, and only the final is checked.

⚠️ script_honoured is checked, not assumed. true means the final is in the alphabet you asked for, false that it is not; null means the final has no letters to check — an empty turn.

How a session ends

Close code Meaning
1000 normal, after session.ended
1009 a frame over 64 KiB
1011 the gateway failed internally
1012 the server is restarting. session.ended is sent first; reconnect, passing the last finals as context
1013 transcription (model_unavailable) or usage recording (billing_unavailable) became unavailable mid-session. Reconnect shortly
4401 key_revoked — the key was revoked during the session
4402 insufficient_credit — the balance ran out, or key_limit_reached — this key's limit did, during the session. The seconds already received are charged
4403 account_suspended — the account was suspended during the session
4408 idle_timeout — no message from you for 60 seconds. Send audio, or ping while a vad=client call is on hold
4413 session_too_long — 3 hours, or 3 hours of audio. Open a new session to continue

The server sends a WebSocket ping every 20 seconds, which your client library answers for you, so a quiet line is not dropped by a proxy on the way. It does not reset the 60-second idle rule, which counts messages from you.

Billed per second of audio you send

₹20/hour — lark-mini's own price — charged on every second of audio the session received, rounded up once per session, not per turn. A 124.37-second call is 125 seconds, 69.4445 paise.

⚠️ Streaming and upload are separate lanes, and their prices may come to differ. Today the live price is the upload price. Hindi live always uses the Indic lane's final pass, because live audio is 8 or 16 kHz by construction — the same rule the upload endpoint applies to telephone-rate Hindi. When the live price changes it changes here and in the changelog first.

⚠️ Silence you send is audio you sent. Ten seconds of a muted line is ten billed seconds with no turn in it. Stop streaming while a call is on hold, and send ping to keep the session.

⚠️ A long session is debited as it runs, not at the end. Every 60 seconds of audio the seconds so far are charged, and your balance and the key's limit are checked again. When the balance is exhausted you get error insufficient_credit, and when the key's limit is, error key_limit_reached — both close 4402, with the seconds received so far charged. A session cannot run past the credit that pays for it.

What this endpoint refuses

Before the upgrade, as the usual JSON error envelope:

Status Code When
400 live_not_available model is anything other than lark-mini
400 language_required no language
400 language_not_supported a language not in the list above
400 script_required an Indian language with no script
400 unknown_script a script that is not latin or native or one of their synonyms — auto included
400 unsupported_encoding an encoding other than linear16, mulaw, alaw
400 unsupported_sample_rate not 8000 or 16000, or 16000 with mulaw/alaw
400 invalid_endpointing endpointing_ms outside 300–1000
400 invalid_vad vad other than server or client
400 domain_too_long / vocabulary_too_long over the limits above
400 invalid_verify_vocabulary verify_vocabulary is not true or false
400 context_too_long more than 3 context values, or one over 500 characters
400 diarize_not_available_live diarize or speakers was sent
401 invalid_api_key / key_revoked as on every endpoint
402 insufficient_credit / key_limit_reached as on every endpoint
403 account_suspended as on every endpoint
426 upgrade_required a plain HTTP request, not a WebSocket upgrade
429 rate_limited live capacity is full right now, or this account already has as many live sessions open as it may. Retry shortly
503 tier_not_deployed live is not enabled for this account, or not serving on this deployment

GET /v1/agent/live — Osprey Live

Live — Hindi and English, on every account (since 0.46.0, 2026-09-17). osprey-live is listed in GET /v1/models with its caller languages.

A voice agent on one WebSocket. You stream the caller's audio in. MiniCrow hears each turn (Lark Live), decides what to say and which of your tools to call (the Osprey Live brain, following your instructions), and speaks the reply (Pica) back down the same socket while it is still being written. You describe the agent once, in session.configure. You run your own tools. MiniCrow runs the three parts and bills each one.

About ₹6.50 per 1 million tokens — means around ₹66 per hour including STT + LLM + TTS (best for AI call agents). An illustration for a typical voice agent with prompt caching, not a quote: the brain is billed per token type, speech-to-text per second of audio and text-to-speech per character.

Languages. The caller's language is the language query parameter.

Language Code Status
Hindi hi Available
English en Available
Marathi mr Coming soon
Tamil ta Coming soon
Telugu te Coming soon
Every other language Coming soon

A language that is coming soon is refused with 400 unsupported_language and a message saying it is coming soon.

wss://api.minicrow.com/v1/agent/live?model=osprey-live&language=hi&encoding=mulaw&sample_rate=8000&output=mulaw_8k&voice=david
Authorization: Bearer mc_...

⚠️ The key goes in the Authorization header of the upgrade, and nowhere else. Not in the URL, and not in a message. A browser's WebSocket cannot set a header, so connect from your server: Osprey Live is for server-side clients, such as the relay between your telephony provider and MiniCrow.

Python

# pip install "websockets>=14"
import asyncio, json, os
from urllib.parse import urlencode

import websockets

params = urlencode({
    "model": "osprey-live", "language": "hi", "encoding": "mulaw", "sample_rate": 8000,
    "output": "mulaw_8k", "voice": "david",
})
URL = f"wss://api.minicrow.com/v1/agent/live?{params}"
HEADERS = {"Authorization": f"Bearer {os.environ['MINICROW_API_KEY']}"}

with open("agent.json", encoding="utf-8") as f:      # your agent spec, see "The agent spec" below
    AGENT = json.load(f)


async def run_tool(ws, call):
    """Your own code: look something up, check a slot, send a message. The socket keeps running meanwhile."""
    if call["name"] == "check_appointment_slots":
        output, failed = {"doctor": call["arguments"]["doctor"], "available": ["10:40", "11:20"]}, False
    else:
        output, failed = {"error": "unknown tool"}, True
    await ws.send(json.dumps({"type": "tool.result", "call_id": call["call_id"], "output": output,
                              "is_error": failed}))


async def send_audio(ws, path):
    with open(path, "rb") as audio:                   # raw 8 kHz mu-law, the caller only, no header
        while chunk := audio.read(160):                # 160 bytes = 20 ms
            await ws.send(chunk)                       # bytes are sent as a binary frame
            await asyncio.sleep(0.02)                  # real-time pace, as a live call arrives
    await ws.send(json.dumps({"type": "end"}))


async def main():
    tasks = []
    try:
        async with websockets.connect(URL, additional_headers=HEADERS) as ws:
            with open("agent-reply.ulaw", "wb") as speaker:
                async for message in ws:               # ends when the server closes
                    if isinstance(message, bytes):     # the agent's voice, in the `output` format
                        speaker.write(message)
                        continue
                    event = json.loads(message)
                    kind = event["type"]
                    if kind == "session.started":
                        await ws.send(json.dumps({"type": "session.configure", "agent": AGENT}))
                    elif kind == "session.configured":
                        for warning in event["warnings"]:
                            print("spec warning:", warning)
                        tasks.append(asyncio.create_task(send_audio(ws, "caller.ulaw")))
                    elif kind == "response.audio.started":
                        print("agent:", event["text"])
                    elif kind == "tool.call":
                        tasks.append(asyncio.create_task(run_tool(ws, event)))
                    elif kind == "turn.usage":
                        print(f"turn {event['turn']}: {event['cost']} paise")
                    elif kind == "session.ended":
                        print(f"session: {event['cost']} paise, reason {event['reason']}")
                    elif kind == "error":
                        print("error:", event["code"], event["message"])
    except websockets.exceptions.InvalidStatus as refused:   # refused before the socket opened
        print(refused.response.status_code, refused.response.body.decode())
    finally:
        for task in tasks:
            task.cancel()


asyncio.run(main())

JavaScript (Node)

// npm install ws — save as agent.mjs, run with node agent.mjs
import fs from "node:fs";
import WebSocket from "ws";

const params = new URLSearchParams({
  model: "osprey-live", language: "hi", encoding: "linear16", sample_rate: "16000", output: "pcm16_24k",
});
const ws = new WebSocket(`wss://api.minicrow.com/v1/agent/live?${params}`, {
  headers: { Authorization: `Bearer ${process.env.MINICROW_API_KEY}` },
});
const agent = JSON.parse(fs.readFileSync("agent.json", "utf8"));
const speaker = fs.createWriteStream("agent-reply.pcm");      // raw 24 kHz 16-bit PCM, no header

// Refused before the socket opened: the body is the usual JSON error.
ws.on("unexpected-response", (req, res) => {
  let body = "";
  res.on("data", (chunk) => (body += chunk));
  res.on("end", () => console.error(res.statusCode, body));
});

async function sendAudio() {
  const audio = fs.readFileSync("caller-16k.pcm");            // raw 16 kHz 16-bit PCM, the caller only
  for (let i = 0; i < audio.length && ws.readyState === WebSocket.OPEN; i += 3200) {   // 3200 bytes = 100 ms
    ws.send(audio.subarray(i, i + 3200));
    await new Promise((resolve) => setTimeout(resolve, 100));
  }
  if (ws.readyState === WebSocket.OPEN) ws.send(JSON.stringify({ type: "end" }));
}

ws.on("message", (data, isBinary) => {
  if (isBinary) return speaker.write(data);                   // the agent's voice
  const event = JSON.parse(data.toString());
  switch (event.type) {
    case "session.started":
      ws.send(JSON.stringify({ type: "session.configure", agent, tool_timeout_ms: 8000 }));
      break;
    case "session.configured":
      sendAudio();
      break;
    case "tool.call": {
      const output = { status: "sent" };                       // run your tool here
      ws.send(JSON.stringify({ type: "tool.result", call_id: event.call_id, output, is_error: false }));
      break;
    }
    case "response.audio.started":
      console.log("agent:", event.text);
      break;
    case "turn.usage":
    case "session.ended":
      console.log(event.type, event.cost, "paise");
      break;
    case "error":
      console.error("error:", event.code, event.message);
      break;
  }
});

ws.on("close", (code, reason) => {
  speaker.end();
  console.log("closed", code, reason.toString());
});

Query

Everything about the audio is decided before the socket opens. A wrong parameter is an ordinary JSON error on the upgrade request, not a socket that opens and then closes.

Parameter Values
model osprey-live required
language required. hi · en the language the caller speaks. Other languages are coming soon: they are refused with unsupported_language, never guessed
encoding required. linear16 · mulaw · alaw the caller's audio. linear16 is signed 16-bit little-endian mono PCM; μ-law and A-law are G.711, as a telephone line carries them
sample_rate required. 8000 · 16000 mulaw and alaw are 8000 only
output mulaw_8k · pcm16_24k the agent's voice. Omitted, it is mulaw_8k when encoding=mulaw and pcm16_24k otherwise. pcm16_24k is raw signed 16-bit little-endian mono PCM at 24 kHz, with no header
voice a built-in voice or one of yours (mcv_…) GET /v1/audio/voices lists both. Omitted, it is david when the agent speaks Hindi and robert when it speaks English
endpointing_ms 300–1000, default 400 how much silence ends the caller's turn. Lower answers sooner and splits more sentences in two
vad server (default) · client client: you decide where the caller's turn ends by sending {"type":"turn_end"}
barge_in draft (default) · speech · off what happens when the caller talks over the agent. See Interruptions
transcripts false (default) · true true sends you each caller turn as input.transcript

Opening a session

  1. The upgrade succeeds, and you receive session.started.
  2. Within 10 seconds, and before any audio, you send session.configure with your agent spec.
  3. You receive session.configured: the spec was accepted. It lists your tools, the limits in force and any warnings about the spec. Start sending audio now.

Audio sent before session.configured is dropped with a non-fatal error not_configured. No session.configure within 10 seconds closes the session with config_timeout (4400).

The agent spec

session.configure carries everything the agent knows: who it is, how it speaks, your instructions, your tools and a few short invented examples. MiniCrow adds its own rules on top of it (how to read a fast draft of the caller's turn, how to behave mid-call, when to use tools) that are the same for every customer. The Osprey Live agent guide covers every field and how to write the parts that decide whether a tool gets called.

{"type":"session.configure",
 "agent":{
   "spec_version":"osprey-live-prompt/0.1",
   "persona":{"name":"Asha"},
   "caller":{"pronoun":"they"},
   "languages":{"caller_speaks":["hi","en"],"agent_speaks":"hi","agent_script":"latin"},
   "speech":{"reprompt":"Ji, boliye?","greeting_words":["hello","namaste"],
             "never_claim_before_result":["booked","sent"],
             "block_before_result":["book ho gaya","sms bhej diya"]},
   "brief":"You are Asha, the appointment desk voice for Leafview Family Clinic, Indore. You are on a live phone call and have already greeted the caller. Speak one short sentence of romanised Hindi per reply. To check or book an appointment, call the tool.",
   "tools":[{"name":"check_appointment_slots","kind":"booking",
             "description":"Check which appointment times are free for a doctor on a date. Commits nothing.",
             "when_to_call":"CALL THIS BEFORE saying any appointment date or time.",
             "when_not_to_call":"Do not call it again for the same doctor and date in the same call.",
             "acknowledgement":"Ji, time dekh leti hoon.",
             "reminder_trigger":"an appointment time",
             "parameters":{"type":"object","properties":{"doctor":{"type":"string"},"date":{"type":"string"}},
                           "required":["doctor","date"]}},
            {"name":"send_sms_confirmation","kind":"action_needs_confirmation",
             "description":"Send an SMS with the appointment details.",
             "when_to_call":"CALL THIS right after the caller says yes to an SMS.",
             "when_not_to_call":"Do not call it again for the same message on the same call.",
             "acknowledgement":"Ji, SMS bhej rahi hoon.",
             "parameters":{"type":"object","properties":{"mobile":{"type":"string"}},"required":["mobile"]},
             "confirm":"runtime"}],
   "examples":[{"history":[{"speaker":"agent","text":"Ji, consultation fee 500 rupees hai."},
                           {"speaker":"customer","text":"Theek hai"}],
                "customer_draft":"कल सुबह दिखाना है",
                "acknowledgement":"Ji, time dekh leti hoon.",
                "tool_call":{"name":"check_appointment_slots","arguments":{"doctor":"Dr. Arjun Rao","date":"2026-03-04"}},
                "tool_result":{"available":["10:40"]},
                "reply":"Kal subah 10:40 par Dr. Rao ka time khali hai ji, 10:40 theek rahega?"}]},
 "tool_timeout_ms":8000,
 "history":[{"speaker":"agent","text":"Namaste, main Asha bol rahi hoon Leafview Family Clinic se."}]}
Field
agent required. The agent spec. spec_version is "osprey-live-prompt/0.1"
tool_timeout_ms 1000–30000, default 8000. How long the agent waits for a tool.result
history at most 64 lines, oldest first, each {"speaker":"agent","text":"…"} or {"speaker":"customer","text":"…"}: what was already said on this call. Use it for a greeting you already played, and when you reconnect

Only in the spec, and never in what the agent is told:

Field
languages.agent_speaks hi or en: the language of the agent's voice. The voice you chose must speak it
speech.block_before_result extra phrases the agent may never say before a tool has returned a result, on top of MiniCrow's own list (see What the runtime enforces). List only past-tense outcome phrases — "book ho gaya", "has been sent" — never ordinary words such as "available"
tools[].confirm model (default) or runtime. See Tools
hearing.vocabulary 0.41.0. Names your callers say that a general listener would not spell right — people, products, places, codes — as the transcript should show them, up to 100. Each entry is a string or {"term": "Dr. Arjun Rao", "spoken": ["doctor arjun rao", "डॉक्टर अर्जुन राव"]} with up to 3 spoken forms. See Hearing vocabulary
tools[].execution {"type":"client"}, the default: you run the tool when you receive tool.call

⚠️ Everything in examples must be invented. The agent copies from examples: names, numbers and dates in them reach real replies. session.configured.warnings flags digit runs that look real and numbers shared with your brief.

Hearing vocabulary

The caller's turn reaches the agent as a fast draft, and a fast draft hears a name it has never seen as the nearest everyday words: "Orbit" as "और भी", "Zenith" as "जैने", "WhatsApp" as "what's that mean". Tell the hearing what to listen for and it is boosted while it decodes — the standard hotwords of speech APIs — at no extra cost and about 20 ms.

"hearing": {
  "vocabulary": [
    "Orbit", "Zenith",
    {"term": "Dr. Arjun Rao", "spoken": ["doctor arjun rao", "डॉक्टर अर्जुन राव"]},
    {"term": "Tarangpur", "spoken": ["तरंगपुर"]}
  ]
}
  • List names, not words. People, products, places, models, codes. An everyday word ("today", "road", "आज") is never boosted, however you write it; boosting common words made the hearing hear them where they were not said.
  • Spell the term as the transcript should show it. That spelling is what the agent reads and what a tool receives.
  • For Hindi and Marathi callers, add the Devanagari form in spoken. The hearing writes those languages in Devanagari, so "Orbit" alone cannot match; {"term": "Orbit", "spoken": ["ऑर्बिट"]} can. English callers need no spoken form. Add a spoken form too when callers say a name differently from how it is written.
  • Your tools' enum values are boosted on their own (skin_specialist is heard as "skin specialist"); do not list them.
  • Limits: 100 entries; a term at most 5 words and 64 characters; 3 spoken forms of at most 64 characters. Over any of these the spec is refused (invalid_agent) with the entry named.
  • Names the vocabulary does not carry — the caller's own town, a person the spec never mentions — are not helped; for those the runtime read-back (confirm: "runtime") is the guard.
  • A yes to the read-back sends the held call (0.43.2): after confirm: "runtime" reads the details back, the caller's "haan", "ho", "बरोबर", "yes" makes the agent's next reply call the tool with the same details; a changed value is read back again. Before 0.43.2 the agent could ask "shall I send it?" a second time instead.
  • A phone number the caller never said is never sent (0.43.1): a placeholder such as 9876543210, or a number that appears only in your examples, in any tool's phone argument is held; the agent is told to use caller_id when the caller said to use the number they are calling from, and to ask for the number otherwise. Measured: the default brain had written 9876543210 — which appears nowhere in the spec — into 6 of 99 tool turns.
  • On for Hindi, Marathi and English — the languages it is measured on. Hindi and Marathi boost at one strength, English at a slightly lower one: on 88 English turns the boosted hearing made fewer word errors than the plain decode (12.4 % against 14.3 %), caught every covered name (13 of 13 against 12) and added one word that was not said. The other 18 languages use the vocabulary as a spelling hint only, without the boost, until each is measured the same way.
  • Hindi and Marathi hearing is locked to Devanagari (0.43.0), with or without a vocabulary. The fast draft is written without knowing the language, and 22 of 166 Hindi and Marathi turns had come back with letters of another Indian script — a name in Bengali letters matches no Devanagari spelling. With the lock: 0 of 166; English words in English letters, digits and punctuation are untouched.

Measured (replay, 166 Hindi and Marathi turns, invented clinic / car service / courier calls plus real Marathi calls): of the spoken names the vocabulary covered, the boosted hearing caught 5 of 7 against 2 of 7 in production, with the word error of the whole set unchanged; with the Devanagari lock as well, 6 of 7 and the word error 0.7 points lower (within the run's noise). The required tool calls of the same brain did not move beyond the replay's run-to-run noise.

Writing a spec that calls its tools

MiniCrow's own rules name no business. Whether the agent calls your tool at the right moment depends on your spec, and session.configured.warnings checks it against these rules. A warning never refuses a spec.

Warning The rule
no when_to_call Every tool says when to call it, starting with "CALL THIS" and naming situations, not your test sentences. A tool described only by what it does was not called
N required fields A lookup or booking with more than 3 required fields waits until the caller has answered all of them. Require only what the tool cannot run without; ask the rest after the result
when_to_call waits for a confirmation A lookup or booking commits nothing, so it should not wait for a yes or for a "confirmed" value. Write "call as soon as <required fields> are known"
is dictated by the caller; set "confirm":"runtime" An action whose arguments include a phone number, name, email, account or order number the caller says aloud. With runtime confirmation the values are read back, numbers digit by digit, and a misheard name or number is asked again or spelt before the call reaches you
should say when NOT to call it An action says when not to call it: already done on this call, a number still being dictated
no example calls it A lookup or booking that no example calls. One invented example that calls it at the right moment teaches more than a rule
no examples A spec without examples made the fewest required calls
names X, which is not one of your tools The brief names a tool that is not in tools, and the agent tries to call it
acknowledgement is over 6 words The spoken line before a tool is 3 to 5 words and promises nothing
looks like a phone or account number See above: examples are invented

Describe each argument as a value and where it comes from ("the phone number to send to: the caller's own number if they say to use it, otherwise the number they give"), not as steps the caller goes through ("the number they dictated and confirmed"): the second wording made the agent ask callers to dictate a number they had just told it to use.

A spec that cannot be used ends the session with error invalid_agent, fatal: true, and close 4400. The message names the field and the reason, for example tools[1].name: say_filler is reserved. It happens when:

  • a required field is missing, a tool name is repeated, a kind is not lookup, booking, action_needs_confirmation or other, or an example calls a tool with arguments its parameters do not allow;
  • the brief is over 24,000 characters, or there are more than 8 tools or more than 12 examples;
  • agent_speaks is not hi or en, or the voice does not speak it;
  • the spec contains a cache_control key anywhere, or an example carries a role;
  • a tool is named say_filler (reserved), or execution is anything but {"type":"client"}.

Keys the format does not use are ignored: any key starting with _, and examples[].id. session.configure may be up to 256 KiB. A second session.configure is a non-fatal already_configured.

What you send

Message
binary frame the caller's audio in the encoding and sample_rate you declared, 20 ms to 1 s per frame. A frame over 64 KiB closes the session with 1009
session.configure once, first. See above
tool.result the answer to a tool.call: {"type":"tool.result","call_id":"call_7f3aQ2mX9kLp","output":{"status":"sent"},"is_error":false}. output is any JSON value or a string, at most 32 KiB
response.say speak this text as it is, with no brain call: {"type":"response.say","text":"Namaste, main Asha bol rahi hoon Leafview Family Clinic se. Kya do minute baat kar sakte hain?"}. At most 500 characters. While a response is running it is refused with a non-fatal busy
response.cancel stop the current response, for example on a keypad press: {"type":"response.cancel"}
response.played optional. How much of a cancelled response the caller really heard: {"type":"response.played","response_id":"r_12","played_ms":1840}
turn_end vad=client only: the caller's turn ends here
end finish. A running response completes (at most 15 seconds), the last usage is charged, then session.ended and close 1000
ping answered with {"type":"pong"}
input.image 0.68.0. a picture from the caller's camera: {"type":"input.image","data":"<base64 JPEG, PNG or WebP>"}, at most 192 KB. Send one every 1–3 seconds while the camera is on. See Vision

Any other text message is answered with a non-fatal error bad_message, and the session carries on.

⚠️ Send only the caller's audio, on one channel. If the agent's own voice comes back in your input (a mixed recording, a speakerphone, a call leg that carries both parties), the agent hears itself talking and stops mid-sentence. Take the caller's inbound track only.

What you receive

Text frames are JSON events. Binary frames are the agent's voice in your output format, at most 200 ms each. One turn with a tool, in order, up to the caller's next turn (which closes turn 7's charge):

{"type":"session.started","session_id":"6f1c2a9e-3b7d-4c1a-9f0e-5d8b7a6c4e21","model":"osprey-live"}
{"type":"session.configured","session_id":"6f1c2a9e-3b7d-4c1a-9f0e-5d8b7a6c4e21","model":"osprey-live","input":{"encoding":"mulaw","sample_rate":8000},"output":{"format":"mulaw_8k","sample_rate":8000},"voice":"david","language":"hi","tools":["check_appointment_slots","send_sms_confirmation"],"limits":{"max_seconds":10800,"tool_timeout_ms":8000,"max_passes":3},"warnings":[]}
{"type":"input.speech_started","turn":7,"start":41.232}
{"type":"input.transcript","turn":7,"stage":"draft","text":"कल सुबह डॉक्टर राव का टाइम है क्या","start":41.232,"end":43.02}
{"type":"response.started","response_id":"r_12","turn":7,"kind":"reply"}
{"type":"response.audio.started","response_id":"r_12","clause":0,"kind":"speech","text":"Ji, time dekh leti hoon.","format":"mulaw_8k"}
{"type":"tool.call","response_id":"r_12","turn":7,"call_id":"call_7f3aQ2mX9kLp","name":"check_appointment_slots","arguments":{"doctor":"Dr. Arjun Rao","date":"2026-03-04"}}
{"type":"response.audio.done","response_id":"r_12","clause":0,"audio_ms":1520}
{"type":"response.audio.started","response_id":"r_12","clause":1,"kind":"speech","text":"Kal subah 10:40 par Dr. Rao ka time khali hai ji, 10:40 theek rahega?","format":"mulaw_8k"}
{"type":"response.audio.done","response_id":"r_12","clause":1,"audio_ms":3880}
{"type":"response.done","response_id":"r_12","turn":7,"status":"completed","passes":2,"tool_calls":1,"sent_ms":5400}
{"type":"input.speech_started","turn":8,"start":52.108}
{"type":"input.transcript","turn":8,"stage":"draft","text":"हां 10:40 ठीक है","start":52.108,"end":53.3}
{"type":"turn.usage","turn":7,"brain":{"input_tokens":412,"cached_input_tokens":4220,"cache_write_tokens":0,"output_tokens":61,"audio_input_tokens":0,"passes":2,"cost":3.3},"voice":{"characters":93,"cost":8.37},"cost":11.67,"cost_currency":"INR_paise","cost_known":true}

Binary frames arrive between each response.audio.started and its response.audio.done. When the agent speaks an acknowledgement before a tool, tool.call arrives right after the acknowledgement's first binary frame, so you can start the tool while the caller hears it.

Event Fields
session.started session_id, model
session.configured session_id, model, input (encoding, sample_rate), output (format, sample_rate), voice, language, tools (names), limits (max_seconds, tool_timeout_ms, max_passes), vision on a session that takes pictures (formats, max_image_bytes, min_interval_ms, frames_per_turn), warnings (always an array)
input.speech_started turn (from 0), start: seconds into the audio you have sent
input.transcript transcripts=true only. turn, stage: "draft", text, start, end. A fast draft, written as it was first heard: usually in the language's own script
response.started response_id, turn, kind: reply (the agent answering a turn) or say (your response.say)
response.audio.started response_id, clause (from 0), kind: speech or filler, text: exactly the words in the audio that follows, format
response.audio.done response_id, clause, audio_ms
tool.call response_id, turn, call_id, name, arguments (an object that matches the tool's parameters)
tool.timeout call_id, after_ms: no result arrived in time, and the agent carried on without it
response.done response_id, turn, status, passes, tool_calls, sent_ms, and reason when status is cancelled. status: completed · cancelled (reason: barge_in or client) · incomplete (the reply stopped part-way) · failed
turn.usage the charge for one turn. See Billing
error code, message, fatal. After a fatal error the session ends: session.ended, then the close
session.ended session_id, reason, turns, totals per part and cost. See Billing
pong —

⚠️ A binary frame belongs to the clause of the latest response.audio.started that has no response.audio.done yet. One writer sends every frame, text and audio, so the order on the socket is the order to play.

⚠️ The audio arrives at the pace it should be played, plus a lead of at most 200 ms. Play it as it arrives; you do not need a large buffer, and when the agent stops (an interruption), it goes quiet within about 200 ms without you clearing anything.

A turn, step by step

  1. The caller speaks; you receive input.speech_started.
  2. The caller stops. After endpointing_ms of silence the turn is drafted.
  3. The brain reads your spec, the call so far and the draft, and starts writing. The first clause goes to the voice at once, and later clauses follow while the brain is still writing. You get response.started, then each clause's response.audio.started, its audio, and response.audio.done.
  4. If the brain calls a tool, it usually speaks your tool's acknowledgement first ("Ji, time dekh leti hoon."), then you receive tool.call. When your tool.result arrives, the brain is asked again with the result, and it speaks the answer. A reply may take up to 3 passes.
  5. response.done. What the caller heard goes into the call's history.
  6. turn.usage arrives once the turn's charge is recorded: after the caller's next turn is drafted, or when the session ends.

The first turns of a call are slower by a few seconds, while your instructions and tools are stored in the prompt cache; later turns read them from it. If no audio has gone out 1.8 seconds after the caller's turn was drafted, or a tool is slow, the agent says a short holding phrase in the same voice ("Ji, ek second."). It arrives as a clause with kind: "filler", it is billed as voice characters, and its words enter the history like any other speech.

A sound under 0.3 seconds (a cough, a "hm") is not a turn and gets no reply. A caller turn is merged with the next one, and not answered alone, when it was cut at 20 seconds of continuous speech, or when the caller started speaking again before the agent's first audio went out. A turn with no clear words gets your speech.reprompt; after two reprompts in a row, a third empty turn gets no reply.

Vision: the caller's camera

0.68.0. Osprey Live sees what the caller's camera sees — smart glasses, a phone, a robot, a kiosk, a car. Send a picture every one to three seconds as input.image, on the same socket as the audio. When the caller finishes a turn, the reply is written with the pictures taken while they spoke — from two seconds before they started, at most three, spread over the turn with the newest always among them — beside their words. So "what is this?" is answered about what the camera showed when it was asked.

{"type":"input.image","data":"/9j/4AAQSkZJRgABAQEASABIAAD…"}
async def camera(ws, frames):  # frames: JPEG bytes from your camera, as they come
    async for jpeg in frames:
        await ws.send(json.dumps({"type": "input.image", "data": base64.b64encode(jpeg).decode()}))
        await asyncio.sleep(1.5)
Formats JPEG, PNG or WebP, base64. A data:image/jpeg;base64,… URL is taken too
Size at most 192 KB a picture; 640 to 1024 pixels wide is plenty
How often one every 1–3 seconds while the camera is on. At most one a second is kept; a faster camera's extra pictures are dropped
Per turn at most 3 pictures go with a reply. A turn shorter than your interval takes the newest picture while it is at most 5 seconds old
What is seen only the turn's own pictures. What the agent said about a picture stays in the call's history as text; earlier pictures are not sent again

A session that takes pictures says so in session.configured: "vision":{"formats":["jpeg","png","webp"],"max_image_bytes":196608,"min_interval_ms":1000,"frames_per_turn":3}. Send pictures after it arrives. A picture that is too large or not JPEG, PNG or WebP — or any picture sent to a session without vision — is answered with a non-fatal error bad_message, and the call goes on.

Billing. Pictures are brain input tokens of the turn they go with: about 1,100 tokens a picture, so a turn with three pictures has about 3,300 more input_tokens in its turn.usage — about 9 paise. A turn without pictures costs what it did.

Tools

  • You run every tool. MiniCrow never calls your systems: it sends tool.call, and you answer with tool.result carrying the same call_id.
  • arguments are checked against the tool's parameters before you see them. A call that does not match is not sent to you; the agent is told what was wrong and carries on.
  • Calls of one pass are sent together, at most 2 per pass, and the agent waits for all of their results.
  • No result within tool_timeout_ms: you receive tool.timeout, and the agent is told the tool timed out. A result that arrives later is still recorded in the history, because the action may have happened, but that reply does not change.
  • output goes to the agent trimmed to 4 KiB, and into the history trimmed to 300 characters. Send short JSON that holds only what the agent may say.
  • is_error: true, a top-level error key, or status failed or error in output count as a failed result.
  • A call_id that is unknown, already answered or timed out gets a non-fatal unknown_call_id. An output over 32 KiB gets result_too_large.

Two confirmation modes, per tool:

confirm
model (default) the brain decides when the caller has agreed, following your when_to_call and when_not_to_call
runtime the first call is not sent to you. The agent is told to read the arguments back to the caller, numbers digit by digit, and to ask the caller to spell or repeat a name, email or number that may have been misheard, then call again with the corrected value. If it calls the same tool with the same arguments within the next 2 caller turns and the caller's turn is a yes ("haan", "ho", "yes", "sahi hai"; not a question such as "bheja?", not a "no"), you receive tool.call. Otherwise the call keeps waiting and the agent is told to answer the caller and read back again. Different arguments replace the waiting call and need their own yes. A phone number that does not have 10 digits (after +91 or a leading 0) is never sent; the agent is told to ask for the number again. Identifiers such as contact_id and 1800/1860 toll-free numbers are not checked. It costs at least one extra pass

Interruptions

The caller talks over the agent when input.speech_started arrives while a response is playing.

barge_in
draft (default) the agent's audio holds at once. If the caller's sound turns out to be under 0.3 seconds (a cough), or a one- or two-word acknowledgement that is not a question ("haan", "ho", "achha", "hmm", "ok", "theek hai"), the held audio resumes where it stopped. Anything else is real speech: the response is cancelled and the caller's words become the next turn. A yes said over a runtime read-back, or over a question the caller has already heard, is kept as the caller's next turn
speech the response is cancelled as soon as the caller starts speaking
off the agent finishes; the caller's turn is answered afterwards

A cancelled response ends with response.done status: "cancelled", reason: "barge_in", and sent_ms. response.cancel does the same with reason: "client". Tool calls already sent stay valid, and their results are still recorded; calls not yet sent are dropped.

Only what the caller heard enters the history. By default, that is every clause that finished playing more than 200 ms before the cancel. If you send response.played within 1 second of the cancel, your played_ms decides instead. A clause the caller heard only in part is left out; it is never cut mid-word.

Opening lines and reconnects

  • Outbound calls (the agent speaks first): send response.say with the greeting after session.configured, or play your own greeting and pass it in history.
  • Inbound calls (the caller speaks first): just start sending audio. The agent answers the first turn.
  • Reconnects: on close 1012, 1013 or a dropped connection, open a new session and send the call so far in history. The agent continues the call and does not greet again.

response.say text is spoken exactly as written, enters the history as the agent's line, and is billed as voice characters.

What the runtime enforces

These hold whatever your instructions, the caller or a tool result say:

  • No claimed outcome without a result. A clause such as "book ho gaya" or "has been sent" is not spoken unless a tool returned a successful result earlier in the same response, or a booking or action_needs_confirmation tool succeeded earlier in the call. MiniCrow's list covers common Hindi (Latin and Devanagari) and English outcome phrases; speech.block_before_result adds yours. A promise of what the agent is about to do ("bhej rahi hoon") is allowed. If nothing is left to say, the agent says your speech.reprompt.
  • No second greeting. After the agent's first line (audio the caller heard counts, even when it was interrupted), a clause that starts with a greeting, in any of the supported scripts, has the greeting removed, and a clause that only re-introduces the agent ("Main Asha bol rahi hoon, …") is dropped. Your speech.reprompt is never changed.
  • No repeated identical action. An action_needs_confirmation call with the same arguments as one that already succeeded on this call is not sent to you again; the agent is told it is already done.
  • No spoken tool talk. A reply never reads out a tool call the agent wrote into its words instead of making it — JSON, a function call or a snake_case name. Those words are taken out before the voice; the rest is spoken.
  • Limits per caller turn: at most 3 passes, at most 2 tool calls per pass, and at most 600 characters spoken per response.

A withheld clause is not spoken, not billed and not put in the history.

Languages and voices

The caller speaks hi (Hindi) or en (English): the language of the session. Other languages are coming soon
The agent speaks hi or en: agent.languages.agent_speaks
Voices any built-in voice or voice of yours that speaks agent_speaks. Defaults david (Hindi) and robert (English)

Say in your brief how the agent writes: "romanised Hindi in Latin letters, never Devanagari", and how it says numbers, prices and times. The voice reads what the agent writes.

Billing

Three parts, each billed on its own, in rupees, from your MiniCrow credits. Prices are fixed when a session opens.

About ₹6.50 per 1 million tokens — means around ₹66 per hour including STT + LLM + TTS (best for AI call agents). An illustration for a typical voice agent with prompt caching, not a quote: the brain is billed per token type, speech-to-text per second of audio and text-to-speech per character.

Part Billed on Price
Hearing (Lark Live) every second of audio the session receives, silence included, rounded up once per session ₹20 per hour
Brain (Osprey Live) the tokens of every pass, by type (below) per million tokens
Voice (Pica) characters spoken to the caller: replies, acknowledgements, holding phrases and response.say ₹9 per 10,000 characters
Brain token type What it is ₹ per million tokens
Input the part of the prompt not read from the cache: the latest lines, the caller's draft, tool results ₹27.50
Cached input the part read from the prompt cache: usually your brief, tools and examples ₹2.75
Cache write storing the fixed part in the cache, charged each time it happens, on top of cached input ₹9.17
Output what the brain writes: the spoken reply, tool calls and its reasoning ₹165.00

audio_input_tokens in the usage events is always 0 on Osprey Live: the caller's audio is billed as hearing.

Every turn reports its charge once it is recorded:

{"type":"turn.usage","turn":4,
 "brain":{"input_tokens":203,"cached_input_tokens":4220,"cache_write_tokens":0,"output_tokens":33,
          "audio_input_tokens":0,"passes":1,"cost":2.2632},
 "voice":{"characters":114,"cost":10.26},
 "cost":12.5232,"cost_currency":"INR_paise","cost_known":true}

That turn read 4,220 tokens from the cache, so the brain cost 2.2632 paise; the 114 spoken characters cost 10.26 paise. The last event gives the totals:

{"type":"session.ended","session_id":"6f1c2a9e-3b7d-4c1a-9f0e-5d8b7a6c4e21","reason":"client_end","turns":12,
 "hearing":{"audio_seconds":184.4,"billed_seconds":185,"cost":102.778},
 "brain":{"input_tokens":2911,"cached_input_tokens":54860,"cache_write_tokens":4220,"output_tokens":602,
          "audio_input_tokens":0,"cost":36.8931},
 "voice":{"characters":1288,"cost":115.92},
 "cost":255.5911,"cost_currency":"INR_paise","cost_known":true}
  • cost is always paise charged, to four decimals. The session's cost is exactly the sum of what was debited: its turns and its hearing. Your dashboard shows the same rows.
  • cached_input_tokens includes tokens stored in the cache during that turn; cache_write_tokens is charged on top of them.
  • passes is how many times the brain was asked in the turn. A tool call usually makes two.
  • ⚠️ On Osprey Live, cost_known: false means part of the usage could not be measured and was not charged, for example when the connection dropped before a pass reported its tokens. cost is then a lower bound, never an estimate. (On other endpoints cost_known keeps the meaning described there.)
  • A cancelled response is still billed for the tokens the brain used and the characters that were sent to the voice.

⚠️ Voice is the largest part of the bill — about half. An illustration from two test calls with a 4,200-token spec, 336 caller turns an hour and about 107 spoken characters a turn: hearing ₹20.00 + brain ₹13.50 + voice ₹32.24 = ₹65.74 per call-hour (about ₹66). Short spoken replies lower it more than anything else. Keep the fixed part (brief, tools, examples) identical for the whole call so it stays cached.

⚠️ The balance is checked while the session runs. Hearing is charged every 60 seconds of audio, and your balance and the key's limit are checked before every turn and every minute. When credit runs out, a running response finishes and is billed, then you get error insufficient_credit (or key_limit_reached), session.ended, and close 4402.

How a session ends

Close code Meaning
1000 normal, after end and session.ended
1009 an audio frame over 64 KiB, a session.configure over 256 KiB, or an input.image over 320 KiB
1011 the gateway failed internally, or the agent could not answer 3 turns in a row (brain_unavailable)
1012 the server is restarting (server_restarting). session.ended is sent first; reconnect with history
1013 hearing, voice or the agent became unavailable mid-session (hearing_unavailable, voice_unavailable, service_unavailable), usage could not be recorded (billing_unavailable), or your client did not read fast enough (slow_consumer). Reconnect shortly with history
4400 invalid_agent (the spec cannot be used) or config_timeout (no session.configure within 10 seconds)
4401 key_revoked: the key was revoked during the session
4402 insufficient_credit or key_limit_reached during the session. What was used so far is charged
4403 account_suspended during the session
4408 idle_timeout: no message or audio from you for 60 seconds
4413 session_too_long: 3 hours. Open a new session with history to continue

Non-fatal errors leave the session open: bad_message, not_configured, already_configured, busy, unknown_call_id, result_too_large, say_too_long, voice_unavailable (part of a reply could not be spoken; it is not billed and not put in the history), brain_unavailable (the agent could not answer this turn and said your reprompt) and brain_interrupted (the reply stopped part-way; what was spoken stays).

The server sends a WebSocket ping every 20 seconds, which your client library answers for you. It does not reset the 60-second idle rule, which counts messages from you.

What this endpoint refuses

Before the upgrade, as the usual JSON error envelope:

Status Code When
400 model_not_found model is not osprey-live
400 unsupported_language no language, or a language other than hi and en. For a language that is coming soon, the message says so
400 invalid_parameter a missing or wrong encoding, sample_rate, output, voice, endpointing_ms, vad, barge_in or transcripts. The error names the param
400 unsupported_parameter script, diarize, domain, vocabulary or context: parameters of the transcription socket that this one does not take
401 invalid_api_key / key_revoked as on every endpoint
402 insufficient_credit / key_limit_reached as on every endpoint
403 account_suspended as on every endpoint
404 not_found Osprey Live is not enabled for this account
426 upgrade_required a plain HTTP request, not a WebSocket upgrade
429 too_many_sessions no Osprey Live session can open right now: this account already has as many open as it may, or the server's sessions are all in use. Retry shortly
429 hearing_busy live hearing capacity is full right now. Retry shortly
503 hearing_unavailable / service_unavailable a part of the agent is not serving right now. Retry shortly

POST /v1/video/summaries — Lark-V

MiniCrow-specific. OpenAI has no shape for "video → timeline", so this one is ours.

curl https://api.minicrow.com/v1/video/summaries \
  -H "Authorization: Bearer mc_..." -F file=@clip.mp4 -F model=lark-v-large
{"model":"lark-v-large",
 "summary":"A colorful test pattern is shown while a voiceover announces a meeting in Pune.",
 "timeline":[{"t":0,"text":"A test pattern displays while a voice says, \"Kal shaam 5:00 baje…\""}],
 "usage":{"prompt_tokens":288,"completion_tokens":183,"cost":9.5278,"cost_currency":"INR_paise"}}

⚠️ lark-v-large is the most accurate tier. It sees the picture and hears the sound together: speech in the clip appears in the timeline, in the script it was spoken in.

⚠️ Billed on what the clip used, not per second — a per-second rate would be a guess dressed as a price. Every response carries its cost.

lark-v-nano and lark-v-mini

These two return the same summary and timeline, and the words spoken with their times too. lark-v-mini is the more accurate of the two. lark-v-nano is coming soon: until it opens it answers 503 tier_not_deployed, and the message names the tier to use. Everything below applies to both:

curl https://api.minicrow.com/v1/video/summaries \
  -H "Authorization: Bearer mc_..." -F file=@clip.mp4 -F model=lark-v-mini -F effort=high
{"model":"lark-v-mini",
 "summary":"Vertical colour bars with a moving diagonal stripe, while a voice repeats a meeting reminder.",
 "timeline":[{"t":0,"text":"Colour-bar test pattern; a voice says \"Kal shaam 5:00 baje…\""}],
 "x_minicrow":{"tier":"lark-v-mini","effort":"high","heard":true,"cost_known":true},
 "usage":{"prompt_tokens":2058,"completion_tokens":1104,"seconds":24,
   "cost":18.983,"cost_currency":"INR_paise"}}
field meaning
x_minicrow.heard whether the clip had a speech track and it was transcribed
transcript the speech track, as heard
transcript_segments the same transcript a sentence at a time, each with when it is said — [{"start":15.04,"end":20.0,"text":"…"}], in seconds; estimated, see below. Absent when nothing in the track was voiced
usage.seconds seconds of speech in the clip; 0 for a silent clip

⚠️ A silent clip is charged nothing for speech and reports heard: false.

⚠️ The whole clip is summarised, however long — never only its first minute.

⚠️ At most 10 minutes of video on these two tiers — 413 video_too_long beyond that.

⚠️ lark-v-nano is the lighter tier, not a cheaper picture. Its price is close to lark-v-mini's, and on accuracy lark-v-mini is currently the stronger of the two. Both tiers are announced separately in GET /v1/models: one of them serving never implies the other does.

⚠️ What is said sits in the timeline where it is said (since 0.62.0), and transcript_segments carries each sentence with its time. The times are an estimate: on one voice without music they land within about a second.

⚠️ If the model does not return parseable JSON you still get its answer, as raw, with timeline_parsed: false. The model has already done the work; throwing the response away would charge you for nothing.

Room to write the answer: output_budget and answer_cut_off

Every Lark-V response reports x_minicrow.output_budget — the most output tokens the answer could use — and x_minicrow.answer_cut_off, true when the answer ran out of it before finishing. Lark-V reasons before it writes, and the reasoning is paid out of the same budget: a 60-second Hindi clip in script=native at effort=high spent 4,094 of its 4,096 tokens thinking and returned an empty answer. Since 0.62.0 the budget is the preset's — 2,048 (mid), 4,096 (high), 8,192 (max) — multiplied by 3 when the summary or the quoted speech is written in an Indian script, by 2 in another script that is not Latin (Arabic, Russian, Japanese, Korean), and by up to 3 more on longer clips on lark-v-nano and lark-v-mini, capped at 32,768. The same 60-second clip now finishes in 4,552 tokens. It is a ceiling, not a charge: you pay for the tokens the answer used. An answer that is still cut off comes back with what it wrote, answer_cut_off: true, and — if the JSON did not survive — raw; script=latin or a lower effort needs less room.

language and script

-F language=ta -F script=native

Declare the language people speak in the video, as on /v1/audio/transcriptions. For the ready languages — mr hi ta te bn gu kn ml en es de fr it ar — it changes four things:

  • How speech is quoted. Every tier knows which language is spoken and quotes it in that language's own script (script=native, the default), romanised (latin), or with no alphabet asked for (auto), never translated.
  • Which lane hears the speech on lark-v-nano and lark-v-mini: the same lane an audio file in that language gets on the matching Lark tier, billed at that lane's rate. x_minicrow.speech_lane says which.
  • How the speech track is transcribed: as /v1/audio/transcriptions transcribes that language, treating the audio as a video soundtrack — only spoken or sung words are written.
  • The language of the answer. The summary and every timeline text (the descriptions of what is seen) are written in the declared language — in its own script for script=native, romanised for latin, in English for en. What people say is still quoted as the first point says; summary_language below picks another language.

x_minicrow.language echoes the language when it was used, and is "" when none was declared or the language is not one of the above — those requests are served exactly as before. An unknown script is 400 unknown_script on every tier.

summary_language

-F language=ta -F summary_language=en

Choose the language of the summary and the timeline descriptions yourself — a Tamil video summarised in English, as above. It overrides the default from language, takes the same script, and accepts any of the 21 live languages: mr hi bn te ta gu kn or ml pa as en es fr de pt it ru ar ja ko (a region such as en-US is ignored). It works on every tier and never changes how speech is heard or quoted.

  • With language too, speech follows language and the answer follows summary_language.
  • On its own, only the answer's language changes; everything else is served as it is with no language. The code-mixed-Hindi rule then applies to quoted speech only, so a hi or mr summary can be written in Devanagari.
  • How faithfully each tier follows a summary_language different from language has not been measured yet.
  • With neither, no answer language is set, exactly as before.

x_minicrow.summary_language echoes the language the answer was asked to be written in ("" when none was chosen). Anything outside the list is 400 unknown_summary_language, before anything is decoded or billed.

⚠️ Not measured yet on real videos: how faithfully each language is followed.

effort — mid · high · max

-F effort=max
preset what you get output budget
mid a handful of entries — the key moments 2,048 tokens
high (default) an entry every few seconds 4,096
max an entry for every distinct moment 8,192

The response echoes the preset in x_minicrow.effort.

An unknown preset is a 400, not a silent default — a caller who typed maximum asked for something specific.

At most 20 MB — a longer video is a job queue, which this endpoint is not. Over that is 413 file_too_large.

Errors specific to this endpoint

status code when
400 unsupported_video the bytes are not a video container, or the container carries no picture
400 empty_video it is a video and it has no measurable length — nothing to summarise
400 unknown_effort a preset that does not exist; nothing was decoded and nothing was billed
400 unknown_script script is not native, latin or auto; nothing was decoded and nothing was billed
400 unknown_summary_language summary_language is not one of the 21 live languages; nothing was decoded and nothing was billed
413 file_too_large over 20 MB
413 video_too_long over 10 minutes, on lark-v-mini
503 tier_not_deployed a tier this deployment does not serve; the message names the one to use instead

GET /v1/models

Public — no key. A client has to be able to see what it may ask for, and the prices are the product.

{"object":"list","data":[
  {"id":"osprey-flash","object":"model","owned_by":"minicrow","label":"Osprey Flash",
   "modes":["speed","intelligence","max","auto"],"supports_tools":true,"available":true},

  {"id":"osprey-flash-lite","object":"model","owned_by":"minicrow","label":"Osprey Flash Lite",
   "modes":["speed","intelligence","auto"],"supports_tools":false,"available":false,
   "unavailable":{"code":"coming_soon",
                  "reason":"Osprey Flash Lite is coming soon. Use osprey-flash or osprey-pro."}}
]}

⚠️ available: false means do not send a request — every call to that id answers 503 tier_not_deployed. unavailable.code is coming_soon for a model that is announced and not open yet, and tier_not_deployed for one that is not serving; reason names what to use instead. Filter on the field rather than on the label — the list changes as tiers open.

⚠️ They are listed rather than hidden, on purpose. A product that is in the price list and absent from the catalogue is one you meet as unknown_model — which reads as a mistake in your code rather than a gap in ours.

osprey-live joined the list when access opened to every account (0.46.0). Its entry carries a languages list, each with a status of ready (today hi and en) or coming_soon; its section is the reference.

The flag is derived from the same function the handlers refuse with, so what this endpoint advertises and what a call actually does cannot drift apart.


GET /healthz

Public. {"ok":true,"service":"minicrow","version":"0.1.4"}. Liveness only — it deliberately does not touch the database, because a probe that did would restart every pod at once on one blip.

/readyz exists but is not published: it reports whether Postgres and Redis are reachable, which is reconnaissance to a stranger and useless to a customer.


Errors

HTTP code When
400 invalid_json the body is not JSON
400 model_required no model field
400 unknown_model no such tier — the message lists the ones that exist
400 unsupported_audio the recording's container is not one MiniCrow can measure, and the charge is per second of it
400 empty_audio the recording is valid and zero seconds long; a model handed silence writes words nobody said
400 script_mismatch an English voice was asked to read Devanagari
400 emotion_not_supported an emotion was sent to a lane that has none
400 lane_not_supported a voice you created was asked for on the expressive lane, which has its own fixed cast
400 attestation_required a clone was sent without attest=true
400 no_reference_clip a multipart create with no file
400 unusable_clip the recording could not be decoded, or is not speech we can use
400 clip_duration the recording is shorter than 3 seconds or longer than 30
400 description_required a JSON create with no description
404 unknown_voice no such voice on this account — the same answer as a typo, on purpose
409 voice_limit_reached this account already holds the maximum number of voices
413 clip_too_large the upload is far larger than a reference clip
429 voice_rate_limited this key has made too many voices in the last hour or day
503 voices_unavailable this deployment serves the built-in voices only
503 tier_not_deployed a tier that is not serving on this deployment
400 unknown_mode the tier exists, the suffix does not. Distinct from unknown_model on purpose: a client switches on the code, and one of those means "stop using this model" while the other means "fix one word"
401 invalid_api_key missing, malformed, unknown or wrong key — always the same answer
401 key_revoked the key was revoked
401 key_expired the key reached the expiry it was created with. Distinct from key_revoked on purpose: nobody stopped this key, its date passed — the fix is to mint a new one, not to find out who revoked it. A key created without an expiry never returns this
402 insufficient_credit prepaid balance at or below zero
429 — rate limited at the edge; see below
503 model_unavailable the model could not be reached. Recorded but not charged — a call that did not happen is not billed. Carries X-MiniCrow-Origin-Status: 502
400 model_refused_request the model refused the request as sent. Check the parameters and the message shape
400 context_too_long the prompt plus max_tokens is longer than the lane can take
429 rate_limited the lane is busy. Try again shortly, or use another mode
503 model_timeout no answer in time. Carries X-MiniCrow-Origin-Status: 504
503 store_unavailable the key store could not be reached
503 embeddings_failed, speech_failed, transcription_failed, video_failed that service could not answer. Carries X-MiniCrow-Origin-Status: 502
404 unknown_endpoint no such path. The message lists the endpoints that exist
405 method_not_allowed the path exists and does not take that method
413 request_too_large the body is past the gateway limit; audio is capped at 25 MB
500 internal_error the gateway itself failed. Nothing was charged
426 upgrade_required a plain HTTP request to the live endpoint, which only speaks WebSocket. The live endpoint's other refusals are in its own table

Why a gateway failure is a 503 and not a 502

⚠️ Nothing from this API ever answers 502, 504 or 52x — and that is not a taste in status codes. Cloudflare sits in front of api.minicrow.com, and it mints those statuses itself for its own edge→origin failures. It cannot tell ours from its own, so it discards our body and serves 16 bytes of error code: 502 in text/plain. Measured on the live host, one status at a time:

origin answered what reached the customer
200, 400, 429, 500, 501, 503, 505, 507, 508, 509, 527, 530 passed through byte-for-byte
502, 504, 520, 521, 522, 523, 524, 525, 526 body discarded, replaced with error code: NNN, text/plain

The zone setting that would stop it — origin_error_page_pass_thru — is Cloudflare Enterprise-only and reads {"value":"off","editable":false} here. So the fix is on our side: a status the edge would swallow never leaves this API. What was a 502 or a 504 is answered as 503 Service Unavailable, with:

  • the same envelope and the same error.code — the field you branch on is unchanged,
  • Retry-After: 5, which a 502 never carried, and
  • X-MiniCrow-Origin-Status, naming the status the gateway itself decided, so a demoted 504 stays distinguishable from a 503 that always was one.

⚠️ Every failure is JSON, including the ones no handler wrote. A mistyped path, a wrong method, a body over the limit and an internal panic all come back as the envelope above. There is no response from this API that a client parsing error.code cannot parse.

Rate limit

50 requests/second sustained, burst 100, per client IP (keyed on Cf-Connecting-Ip). This is far above what a real caller generates — a chat call takes seconds — because it does not exist to ration the product. It exists to bound the cost of refusing: every wrong key costs one argon2id hash at 19 MiB before it can be rejected.


What a failure will not tell you

⚠️ A failure is described in MiniCrow's words. The message and the code say what you can act on — context_too_long and rate_limited mean exactly what they say.

⚠️ A 401, 402 or 404 is always about your own key, account or request. A lane that fails for a reason of ours is answered as model_unavailable. If it came back as 401 you would check your key; as 402 you would top up. Neither would help: those are our problems with a lane, not yours with your account.

⚠️ Codes changed on 2026-09-11. Four earlier 5xx codes were replaced by the table above. Done while the only client was one we control; a code is a contract, and waiting would have cost more.

What is not here yet

Lark Live preview, enabled per account — every other account is refused 503 tier_not_deployed
Osprey Flash Lite coming soon — refused 503 tier_not_deployed; the message names osprey-flash and osprey-pro
Lark Nano coming soon — refused 503 tier_not_deployed; the message names lark-mini
Lark-V Nano coming soon — refused 503 tier_not_deployed; the message names lark-v-mini

Everything else in the catalogue serves. GET /v1/models is the live answer; this table is a summary of it and can only ever be staler.

When a mode is unavailable but the model is not

GET /v1/models lists the modes a model accepts, and names any that refuse. The example below is the shape; today every listed mode serves, lark-large:max included:

{ "id": "lark-large", "available": true, "modes": ["max"],
  "unavailable_modes": [{ "mode": "max", "code": "tier_not_deployed",
    "reason": "This mode is announced but not wired up yet. Leave the mode off and Lark Large will serve." }] }

⚠️ A TIER THAT SERVES CAN CARRY A MODE THAT DOES NOT, AND THE CATALOGUE USED TO SAY NOTHING. lark-large answers and lark-large:max returned 503, so a customer read available: true and got a refusal from a mode this very response advertised. Found by calling every mode of every listed model, not by reading the list.

⚠️ lark-large TAKES ONE MODE, max, AND NOTHING ELSE. default and indic are chosen for you from the language you declare — asking for them is 400 unknown_mode, which is correct: they are how MiniCrow describes the lane it picked, not something to request.