< Developer Docs · v1 >

Build on MiniCrow

Add Indian-language AI to your own backend: speech to text with speaker labels, text to speech, a real-time voice agent for phone calls and LLM chat. One OpenAI-compatible base URL, Authorization: Bearer mc_…, and the cost of every call in its own response.

Overview

What you get

MiniCrow is one API for the AI a voice product in India needs. You choose the product per request — a model id in the body or a WebSocket to connect to — and every answer says what it cost.

  • One base URL, OpenAI's shape

    POST /v1/chat/completions, /v1/audio/transcriptions, /v1/audio/speech, /v1/video/summaries, /v1/embeddings and two WebSockets for live transcription and the voice agent. An existing OpenAI client works by changing the base URL and the key.

  • Every response says which lane answered, and why

    x_minicrow.route_reason names the rule that picked the lane on every chat call — a router that hides its choice cannot be debugged.

  • Prepaid credit, cost in the response

    Credit lives on your account in rupees; every key draws on it. usage.cost is the charge in paise. At zero, requests stop with 402 before any model is called — see pricing.

Production API base

https://api.minicrow.com

OpenAI Chat Completions-compatible. Verified against the deployed service.

Dashboard

https://dash.minicrow.com

API keys, credit, usage and every request's cost live here.

Quickstart

Make your first call

  1. Create an account at dash.minicrow.com and confirm your email — confirming the address is what unlocks making a key.
  2. Add credit under Billing. Keys are prepaid in rupees.
  3. Create an API key (mc_…). It is shown once — store it in a server-side environment variable such as MINICROW_API_KEY.
  4. Call POST https://api.minicrow.com/v1/chat/completions with Authorization: Bearer your key.
curl https://api.minicrow.com/v1/chat/completions \
  -H "Authorization: Bearer $MINICROW_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "osprey-flash",
    "messages": [{ "role": "user", "content": "Aaj ka plan kya hai?" }]
  }'

Authentication

API keys

Every endpoint takes Authorization: Bearer mc_<prefix>_<secret>. A key looks like mc_96bf0550e045_kQ3v…: the middle segment is the prefix, which identifies the key in a listing and is not secret. The whole string is the secret and is shown once, at creation — only a hash is stored, so a lost key is replaced, not recovered.

Every authentication failure answers the same 401 invalid_api_key, whether the key does not exist or the secret is wrong. A revoked key answers 401 key_revoked. Never auto-retry a 401.

On the two WebSocket endpoints the key goes in the Authorization header of the upgrade — never in the URL and never in a message. A browser WebSocket cannot set that header, so connect from your server.

Billing

Credits, keys and spend limits

Credit belongs to the account; keys draw on it. Make one key per service or environment — they all spend the same balance. A key can carry its own spend limit, so a script or a staging environment stops at its own cap while the rest of your balance stays usable.

StatusMeaning
402 insufficient_creditThe account is out of credit. Top up in the dashboard.
402 key_limit_reachedThis key hit its own limit; the account has not.
403 account_suspendedThe account is suspended. Rotating the key will not help.

A key's limit is a lifetime limit — nothing resets it, so a ₹500 cap is ₹500 for the life of the key. A key with no limit set has no limit. Both 402s are checked before any model is called, so neither costs anything.

Models

Models and modes

The model id selects the product and tier. GET /v1/models is public, needs no key, and lists every tier with available — filter on that field: an id with available: false answers 503 tier_not_deployed.

ModelWhat it doesEndpoint
osprey-flash · osprey-proLLM chat, tools and vision · osprey-flash-lite coming soon/v1/chat/completions
lark-mini · lark-largeSpeech to text from a recording · lark-nano coming soon/v1/audio/transcriptions
lark-mini (live)Streaming speech to text, turn by turn/v1/audio/transcriptions/live
pica-nano · pica-small · pica-largeText to speech, streaming, per second of audio/v1/audio/speech
osprey-liveReal-time voice agent — live/v1/agent/live
lark-v-mini · lark-v-largeVideo → summary and timeline · lark-v-nano coming soon/v1/video/summaries
minicrow-embed · minicrow-rerankEmbeddings (dense + sparse) and rerank/v1/embeddings · /v1/rerank

The chat model id carries the mode

You sendYou get
osprey-flashThe tier's default lane, reported as default:<mode>
osprey-flash:speedThat lane, always — an explicit mode is a hard choice
osprey-flash:intelligenceThat lane, always
osprey-flash:maxThat lane, always
osprey-flash:autoThe router picks, and says why in route_reason

Osprey

Chat completions (POST /v1/chat/completions)

The standard OpenAI request body — messages, tools, images — on MiniCrow's lanes. Two things are added to every response and nothing is taken away: the lane that answered and why, and the cost.

from openai import OpenAI

client = OpenAI(base_url="https://api.minicrow.com/v1", api_key="mc_YOUR_KEY")

r = client.chat.completions.create(
    model="osprey-flash:auto",
    messages=[{"role": "user", "content": "Kal ki meeting ka agenda likho"}],
)
print(r.choices[0].message.content)
print(r.usage.cost, "paise")

Every response says which lane answered

"x_minicrow": {
  "requested_mode": "auto",
  "served_mode":    "intelligence",
  "route_reason":   "S:tools",
  "cost_known":     true
}

route_reason distinguishes explicit:<mode> (you named it), default:<mode> (a bare model id) and the router's own rules such as S:tools. Every MiniCrow model presents itself as MiniCrow's and will not name the model behind a lane.

Streaming

Server-sent events

"stream": true returns text/event-stream, one data: {…} per event, flushed as it arrives and terminated by data: [DONE]. The final chunk carries usage.cost and x_minicrow. Omit stream for one JSON response.

A failure before the first event is the ordinary JSON error, not a stream, and is not charged. A stream that breaks after it started ends on an error event with no [DONE] after it, so an OpenAI client raises instead of treating the partial answer as whole:

data: {"error":{"message":"The answer was cut off before it was complete. Try again.","type":"api_error","code":"model_unavailable"}}

A stream that breaks part-way is still charged for what the model produced, and hanging up does not stop the charge: the model bills the whole generation, which arrives in the final usage chunk.

Lark

Speech to text (POST /v1/audio/transcriptions)

multipart/form-data, OpenAI's shape. Send the recording, the tier, and the language — MiniCrow never guesses it, and the declaration picks the right lane.

curl https://api.minicrow.com/v1/audio/transcriptions \
  -H "Authorization: Bearer mc_YOUR_KEY" \
  -F file=@call.wav -F model=lark-mini -F language=hi -F script=latin

Response

{
  "text": "Kal shaam 5:00 baje Pune me meeting hai.",
  "model": "lark-mini",
  "x_minicrow": { "tier": "lark-mini", "cost_known": true },
  "usage": {
    "seconds": 2.05,
    "cost": 0.3928,
    "cost_currency": "INR_paise",
    "duration_estimated": false,
    "audio_format": "wav"
  }
}
TierRateNotes
lark-nano₹10 / hourComing soon · at most 5 minutes of audio per request
lark-mini₹20 / hourDefault. Hindi and English follow your recording's sample rate
lark-large₹30 / hour · ₹45 for Indian languagesPremium lane
lark-large:max₹60 / hourTop lane for both language groups

Billed per second of audio, read from the file's own header — exact for WAV and OGG/Opus, estimated for MP3, M4A and WebM (duration_estimated says which). At most 25 MB per request. FLAC and raw PCM are refused with 400 unsupported_audio.

Languages

Languages and scripts

21 languages: hi, the Indian languages mr bn te ta gu kn or ml pa as, and en es fr de pt it ru ar ja ko. For an Indian language, say which alphabet you want back:

-F script=native     # each language in its own script — code-mixed English stays in English
-F script=latin      # romanised — "kal shaam ko meeting hai"
-F script=auto       # whatever Lark writes, no preference either way

native keeps code-mixing as spoken: a Marathi speaker who says meeting gets उद्या meeting आहे का?, never a transliteration. The response reports x_minicrow.script_honoured, and script_repaired: true when a transcript came back in the wrong alphabet and was converted — at no charge.

Diarization

Speaker labels — who spoke when

Add diarize=true (and speakers=2 when you know the count) to get a speaker number with start and end times for every turn. +₹3.50 per hour of audio on top of the tier.

curl https://api.minicrow.com/v1/audio/transcriptions \
  -H "Authorization: Bearer mc_YOUR_KEY" \
  -F file=@call.wav -F model=lark-mini -F language=mr \
  -F diarize=true -F speakers=2

Response

{
  "text": "Hello sir, mi Pooja bolte Samarth Sky project madhun. Ha bola, kay aahe sanga...",
  "segments": [
    { "start": 0.0, "end": 4.2, "speaker": 1 },
    { "start": 4.5, "end": 7.1, "speaker": 2 }
  ],
  "x_minicrow": { "speakers": 2 }
}
  • The transcript is unchanged — labels are not interleaved into text; join them on the times.
  • If the speaker pass fails you still get the transcript, and the ₹3.50 is not charged.
  • A two-channel WAV with one party per channel is transcribed channel by channel, billed once by the length of the call — and diarize is then skipped and not charged.

Lark Live

Live speech to text over a WebSocket

Preview, enabled per account. On any other account the upgrade answers 503 tier_not_deployed. Ask us to enable yours.

Stream a call's audio as it happens. Every turn is answered twice: a fast draft the moment the turn ends, then a checked final. Everything is decided in the query, before the socket opens.

wss://api.minicrow.com/v1/audio/transcriptions/live?model=lark-mini&language=mr&script=latin&encoding=mulaw&sample_rate=8000
Authorization: Bearer mc_YOUR_KEY
ParameterValues
languageRequired. Any of the 21 languages
scriptlatin or native — required for every Indian language, Hindi included
encodinglinear16 (default) · mulaw · alaw
sample_rate8000 · 16000 (mulaw and alaw are 8000 only)
endpointing_ms300–1000, default 400 — how much silence ends a turn
vadserver (default) · client — you send {"type":"turn_end"}

What you receive

{"type":"session.started","session_id":"6f1c…","model":"lark-mini","lane":"indic","language":"mr","script":"latin"}
{"type":"speech.started","turn":0,"start":0.42}
{"type":"transcript.partial","turn":0,"stage":"draft","text":"हो सर फ्लॅट चा बुकिंग अमाउंट पन्नास हजार आहे","start":0.42,"end":3.18}
{"type":"transcript.final","turn":0,"text":"Ho sir, flat cha booking amount 50 hazar aahe.","start":0.42,"end":3.18,"script_honoured":true}
{"type":"session.ended","audio_seconds":124.37,"turns":31,"billed_seconds":125,"cost":69.4445,"cost_currency":"INR_paise"}

Send audio as binary frames (20 ms to 1 s each) and {"type":"end"} when done — do not just hang up, or the last finals are lost. ₹20 / hour, billed per second received, rounded up once per session.

Pica

Text to speech (POST /v1/audio/speech)

Returns audio/wav. The charge rides in the response headers, so you can read what a synthesis cost without parsing the audio. Billed per character of input, at most 5,000 characters per request.

curl https://api.minicrow.com/v1/audio/speech \
  -H "Authorization: Bearer mc_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "input": "नमस्ते, आज मौसम बहुत अच्छा है।", "voice": "david", "mode": "natural" }' \
  --output hello.wav

Response headers

X-Voice: david          X-Mode: natural        X-Lane: standard
X-Cost-Paise: 10.8000   X-Cost-Currency: INR_paise
X-Duration-S: 2.05      X-Sample-Rate: 24000
FieldNotes
voiceA built-in voice or one of yours (mcv_…). Omitted: david for Devanagari text, robert otherwise
modenormal or natural (default) — delivery style
lanestandard (default, Hindi and English) or expressive (Hindi, with emotions)
emotionexpressive lane only: neutral · happy · sad · angry · disgust · fear · surprise
streamtrue streams length-prefixed WAV frames, one per sentence
first_clausetrue speaks the first clause on its own — first audio about 437 ms sooner at p50
response_formatOn a stream, "mulaw_8k" sends 8 kHz G.711 μ-law frames for a phone line

Streaming to a phone line

{
  "input": "आज शाम चार बजे, राजेश के साथ आपकी कॉल तय है।",
  "voice": "david",
  "stream": true,
  "first_clause": true,
  "response_format": "mulaw_8k"
}

emotion on the standard lane is a 400, not an ignored field, and an English voice will not read Devanagari (400 script_mismatch) — a flat or nonsense reading billed as if it worked is worse than a refusal.

Voices

List voices (GET /v1/audio/voices)

Eleven built-in voices — seven Hindi, four English — each pinned to a fixed reference clip, so the voice you shipped last month is the voice you get today; plus any voices your account has made.

{
  "object": "list", "built_in_count": 11, "own_count": 1,
  "data": [
    { "id": "david", "name": "David", "language": "hi", "gender": "male", "built_in": true },
    { "id": "mcv_k4t2…", "name": "Anil", "language": "hi", "kind": "generated", "built_in": false }
  ],
  "modes": ["normal", "natural"],
  "lanes": ["standard", "expressive"]
}

Osprey Live

Real-time voice agent over one WebSocket

Live — Hindi and English, on every account. Marathi, Tamil, Telugu and more are coming soon.

You stream the caller's audio in. MiniCrow hears each turn, decides what to say and which of your tools to call, and speaks the reply back down the same socket while it is still being written. About ₹66 per call-hour for hearing, brain and voice together, as an illustration.

wss://api.minicrow.com/v1/agent/live?model=osprey-live&language=hi&encoding=mulaw&sample_rate=8000&output=mulaw_8k&voice=david
Authorization: Bearer mc_YOUR_KEY
  1. The upgrade succeeds and you receive session.started.
  2. Within 10 seconds, before any audio, send session.configure with the agent spec.
  3. On session.configured, start sending the caller's audio as binary frames — the caller's track only, on one channel.

session.configure

{
  "type": "session.configure",
  "agent": {
    "spec_version": "osprey-live-prompt/0.1",
    "persona": { "name": "Asha" },
    "languages": { "caller_speaks": ["hi", "en"], "agent_speaks": "hi", "agent_script": "latin" },
    "brief": "You are Asha, the appointment desk voice for Leafview Family Clinic. Speak one short sentence of romanised Hindi per reply. To check or book an appointment, call the tool.",
    "tools": [{ "name": "check_appointment_slots", "kind": "booking" }]
  }
}

One turn with a tool

{"type":"input.speech_started","turn":7,"start":41.232}
{"type":"response.started","response_id":"r_12","turn":7,"kind":"reply"}
{"type":"response.audio.started","response_id":"r_12","clause":0,"kind":"speech","text":"Ji, time dekh leti hoon."}
{"type":"tool.call","response_id":"r_12","call_id":"call_7f3aQ2mX9kLp","name":"check_appointment_slots","arguments":{"doctor":"Dr. Arjun Rao","date":"2026-03-04"}}
{"type":"response.done","response_id":"r_12","turn":7,"status":"completed","passes":2,"tool_calls":1}
{"type":"turn.usage","turn":7,"brain":{"cost":3.3},"voice":{"characters":93,"cost":8.37},"cost":11.67,"cost_currency":"INR_paise"}

Answer a tool.call

{ "type": "tool.result", "call_id": "call_7f3aQ2mX9kLp", "output": { "status": "sent" }, "is_error": false }

Binary frames between response.audio.started and response.audio.done are the agent's voice, paced for playback. Send response.say to speak fixed text, response.cancel to stop, and end to finish. Every field is in the full reference.

Lark-V

Video summaries (POST /v1/video/summaries)

One upload returns a summary and a seekable, timestamped timeline. lark-v-large watches the video natively — it sees the frames and hears the speech. At most 20 MB.

curl https://api.minicrow.com/v1/video/summaries \
  -H "Authorization: Bearer mc_YOUR_KEY" \
  -F file=@clip.mp4 -F model=lark-v-large -F effort=high

Response

{
  "model": "lark-v-large",
  "summary": "A colorful test pattern is shown while a voiceover announces a meeting in Pune.",
  "timeline": [{ "t": 0, "text": "A test pattern displays while a voice says, \"Kal shaam 5:00 baje…\"" }],
  "usage": { "prompt_tokens": 288, "completion_tokens": 183, "cost": 9.5278, "cost_currency": "INR_paise" }
}
effortWhat you get
midThe key moments
high (default)A moment every few seconds
maxEvery distinct moment

Retrieval

Embeddings and rerank

POST /v1/embeddings returns dense and sparse vectors from one call — the sparse half is what hybrid search needs. At most 256 inputs per request. POST /v1/rerank scores a shortlist against a query and is billed on query plus documents.

/v1/embeddings

{ "model": "minicrow-embed", "input": ["kal meeting hai Pune me", "tomorrow there is a meeting"] }

/v1/rerank

{
  "model": "minicrow-rerank",
  "query": "Pune meeting kab hai",
  "documents": ["Delhi ka flight subah 6 baje", "Pune me meeting kal 3 baje hai"],
  "top_n": 2
}

Response shape

The cost of every call, in paise

usage.cost is always MiniCrow's charge, decimal, in INR_paise — the same number in a stream as without one. A short call costs a fraction of a paisa, which is why the field is not rounded.

"usage": {
  "prompt_tokens": 88,
  "completion_tokens": 60,
  "cost": 0.6458,
  "cost_currency": "INR_paise"
}
EndpointWhere the cost is
Chat, transcription, video, embeddingsusage.cost in the JSON body
Text to speechX-Cost-Paise response header
Live transcriptioncost on session.ended
Voice agentcost on every turn.usage

cost_known: false means no price was reported and the lane has no flat rate — nothing was charged. It is a gap we show you, not a discount.

Errors

HTTP semantics and what you should do

Every failure is the same JSON envelope — {"error":{"message":"…","type":"…","code":"…"}} — so a client can always branch on error.code.

HTTPcodeWhenWhat to do
200—SuccessRead usage.cost
400unknown_model / unknown_modeNo such tier, or a mode that tier does not offerFix the model id; the message lists valid ones
400context_too_longPrompt plus max_tokens exceeds the laneTrim the input or lower max_tokens
400unsupported_audio / empty_audioA container we cannot measure, or zero secondsSend WAV, OGG/Opus, MP3, M4A or WebM
401invalid_api_key / key_revokedMissing, wrong or revoked keyStop and alert an operator; do not retry
402insufficient_creditBalance at or below zeroTop up, then retry
402key_limit_reachedThis key hit its own limitRaise the cap or use another key
403account_suspendedThe account is suspendedContact support
413file_too_largeAudio over 25 MB, video over 20 MBSplit the file
429rate_limitedThe lane is busy, or too many requestsBack off and retry, or use another mode
503model_unavailable / model_timeoutThe model could not answer in timeRetry after Retry-After; not charged
503tier_not_deployedThe tier or mode is not servingUse the tier the message names

A model provider's own 401, 402 or 404 is answered as model_unavailable, never as your credential or credit problem. The full reference lists every code.

Rate limits

Limits, retries and 503s

50 requests per second sustained, burst 100, per client IP. That is far above what a real integration sends; it exists to bound the cost of refusing bad keys. A busy lane answers 429 rate_limited — back off exponentially (1 s → 2 s → 4 s, capped around 30 s) and retry.

MiniCrow never answers 502 or 504: a gateway failure comes back as 503 with the same error.code, a Retry-After: 5 header, and X-MiniCrow-Origin-Status naming what the gateway decided. model_unavailable is recorded but not charged.

For AI coding agents

Integrate MiniCrow using Cursor, Antigravity or your own agent

We publish a single, opinionated agent.md that lays out everything a coding agent needs to integrate MiniCrow end to end: which model for which job, the request shapes, streaming, and the language and script rules that decide accuracy. Drop it at the root of your project — modern agents auto-load AGENTS.md from there.

Auto-detected by

Cursor & Cursor CLIGoogle AntigravityGitHub Copilot WorkspaceContinue.devOpenAI Codex CLIAider (--read AGENTS.md)

Two-second install (run in your project root)

curl -fsSL https://www.minicrow.com/agent.md -o AGENTS.md
git add AGENTS.md && git commit -m "docs: add MiniCrow AGENTS.md"

Then prompt your agent: “Integrate MiniCrow into this app — read AGENTS.md, store the API key as MINICROW_API_KEY, and add a transcription endpoint using lark-mini.” Agents follow the file end to end without further context.