< Developer Docs · v1 >
Build on MiniCrow
Add Indian-language AI to your own backend: speech to text with speaker labels, text to speech, a real-time voice agent for phone calls and LLM chat. One OpenAI-compatible base URL, Authorization: Bearer mc_…, and the cost of every call in its own response.
Overview
What you get
MiniCrow is one API for the AI a voice product in India needs. You choose the product per request — a model id in the body or a WebSocket to connect to — and every answer says what it cost.
One base URL, OpenAI's shape
POST /v1/chat/completions,/v1/audio/transcriptions,/v1/audio/speech,/v1/video/summaries,/v1/embeddingsand two WebSockets for live transcription and the voice agent. An existing OpenAI client works by changing the base URL and the key.Every response says which lane answered, and why
x_minicrow.route_reasonnames the rule that picked the lane on every chat call — a router that hides its choice cannot be debugged.Prepaid credit, cost in the response
Credit lives on your account in rupees; every key draws on it.
usage.costis the charge in paise. At zero, requests stop with402before any model is called — see pricing.
Production API base
https://api.minicrow.com
OpenAI Chat Completions-compatible. Verified against the deployed service.
Dashboard
https://dash.minicrow.com
API keys, credit, usage and every request's cost live here.
Quickstart
Make your first call
- Create an account at dash.minicrow.com and confirm your email — confirming the address is what unlocks making a key.
- Add credit under Billing. Keys are prepaid in rupees.
- Create an API key (
mc_…). It is shown once — store it in a server-side environment variable such asMINICROW_API_KEY. - Call
POST https://api.minicrow.com/v1/chat/completionswithAuthorization: Beareryour key.
curl https://api.minicrow.com/v1/chat/completions \
-H "Authorization: Bearer $MINICROW_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "osprey-flash",
"messages": [{ "role": "user", "content": "Aaj ka plan kya hai?" }]
}'Authentication
API keys
Every endpoint takes Authorization: Bearer mc_<prefix>_<secret>. A key looks like mc_96bf0550e045_kQ3v…: the middle segment is the prefix, which identifies the key in a listing and is not secret. The whole string is the secret and is shown once, at creation — only a hash is stored, so a lost key is replaced, not recovered.
Every authentication failure answers the same 401 invalid_api_key, whether the key does not exist or the secret is wrong. A revoked key answers 401 key_revoked. Never auto-retry a 401.
On the two WebSocket endpoints the key goes in the Authorization header of the upgrade — never in the URL and never in a message. A browser WebSocket cannot set that header, so connect from your server.
Billing
Credits, keys and spend limits
Credit belongs to the account; keys draw on it. Make one key per service or environment — they all spend the same balance. A key can carry its own spend limit, so a script or a staging environment stops at its own cap while the rest of your balance stays usable.
| Status | Meaning |
|---|---|
| 402 insufficient_credit | The account is out of credit. Top up in the dashboard. |
| 402 key_limit_reached | This key hit its own limit; the account has not. |
| 403 account_suspended | The account is suspended. Rotating the key will not help. |
A key's limit is a lifetime limit — nothing resets it, so a ₹500 cap is ₹500 for the life of the key. A key with no limit set has no limit. Both 402s are checked before any model is called, so neither costs anything.
Models
Models and modes
The model id selects the product and tier. GET /v1/models is public, needs no key, and lists every tier with available — filter on that field: an id with available: false answers 503 tier_not_deployed.
| Model | What it does | Endpoint |
|---|---|---|
| osprey-flash · osprey-pro | LLM chat, tools and vision · osprey-flash-lite coming soon | /v1/chat/completions |
| lark-mini · lark-large | Speech to text from a recording · lark-nano coming soon | /v1/audio/transcriptions |
| lark-mini (live) | Streaming speech to text, turn by turn | /v1/audio/transcriptions/live |
| pica-nano · pica-small · pica-large | Text to speech, streaming, per second of audio | /v1/audio/speech |
| osprey-live | Real-time voice agent — live | /v1/agent/live |
| lark-v-mini · lark-v-large | Video → summary and timeline · lark-v-nano coming soon | /v1/video/summaries |
| minicrow-embed · minicrow-rerank | Embeddings (dense + sparse) and rerank | /v1/embeddings · /v1/rerank |
The chat model id carries the mode
| You send | You get |
|---|---|
| osprey-flash | The tier's default lane, reported as default:<mode> |
| osprey-flash:speed | That lane, always — an explicit mode is a hard choice |
| osprey-flash:intelligence | That lane, always |
| osprey-flash:max | That lane, always |
| osprey-flash:auto | The router picks, and says why in route_reason |
Osprey
Chat completions (POST /v1/chat/completions)
The standard OpenAI request body — messages, tools, images — on MiniCrow's lanes. Two things are added to every response and nothing is taken away: the lane that answered and why, and the cost.
from openai import OpenAI
client = OpenAI(base_url="https://api.minicrow.com/v1", api_key="mc_YOUR_KEY")
r = client.chat.completions.create(
model="osprey-flash:auto",
messages=[{"role": "user", "content": "Kal ki meeting ka agenda likho"}],
)
print(r.choices[0].message.content)
print(r.usage.cost, "paise")Every response says which lane answered
"x_minicrow": {
"requested_mode": "auto",
"served_mode": "intelligence",
"route_reason": "S:tools",
"cost_known": true
}route_reason distinguishes explicit:<mode> (you named it), default:<mode> (a bare model id) and the router's own rules such as S:tools. Every MiniCrow model presents itself as MiniCrow's and will not name the model behind a lane.
Streaming
Server-sent events
"stream": true returns text/event-stream, one data: {…} per event, flushed as it arrives and terminated by data: [DONE]. The final chunk carries usage.cost and x_minicrow. Omit stream for one JSON response.
A failure before the first event is the ordinary JSON error, not a stream, and is not charged. A stream that breaks after it started ends on an error event with no [DONE] after it, so an OpenAI client raises instead of treating the partial answer as whole:
data: {"error":{"message":"The answer was cut off before it was complete. Try again.","type":"api_error","code":"model_unavailable"}}A stream that breaks part-way is still charged for what the model produced, and hanging up does not stop the charge: the model bills the whole generation, which arrives in the final usage chunk.
Lark
Speech to text (POST /v1/audio/transcriptions)
multipart/form-data, OpenAI's shape. Send the recording, the tier, and the language — MiniCrow never guesses it, and the declaration picks the right lane.
curl https://api.minicrow.com/v1/audio/transcriptions \
-H "Authorization: Bearer mc_YOUR_KEY" \
-F file=@call.wav -F model=lark-mini -F language=hi -F script=latinResponse
{
"text": "Kal shaam 5:00 baje Pune me meeting hai.",
"model": "lark-mini",
"x_minicrow": { "tier": "lark-mini", "cost_known": true },
"usage": {
"seconds": 2.05,
"cost": 0.3928,
"cost_currency": "INR_paise",
"duration_estimated": false,
"audio_format": "wav"
}
}| Tier | Rate | Notes |
|---|---|---|
| lark-nano | ₹10 / hour | Coming soon · at most 5 minutes of audio per request |
| lark-mini | ₹20 / hour | Default. Hindi and English follow your recording's sample rate |
| lark-large | ₹30 / hour · ₹45 for Indian languages | Premium lane |
| lark-large:max | ₹60 / hour | Top lane for both language groups |
Billed per second of audio, read from the file's own header — exact for WAV and OGG/Opus, estimated for MP3, M4A and WebM (duration_estimated says which). At most 25 MB per request. FLAC and raw PCM are refused with 400 unsupported_audio.
Languages
Languages and scripts
21 languages: hi, the Indian languages mr bn te ta gu kn or ml pa as, and en es fr de pt it ru ar ja ko. For an Indian language, say which alphabet you want back:
-F script=native # each language in its own script — code-mixed English stays in English
-F script=latin # romanised — "kal shaam ko meeting hai"
-F script=auto # whatever Lark writes, no preference either waynative keeps code-mixing as spoken: a Marathi speaker who says meeting gets उद्या meeting आहे का?, never a transliteration. The response reports x_minicrow.script_honoured, and script_repaired: true when a transcript came back in the wrong alphabet and was converted — at no charge.
Diarization
Speaker labels — who spoke when
Add diarize=true (and speakers=2 when you know the count) to get a speaker number with start and end times for every turn. +₹3.50 per hour of audio on top of the tier.
curl https://api.minicrow.com/v1/audio/transcriptions \
-H "Authorization: Bearer mc_YOUR_KEY" \
-F file=@call.wav -F model=lark-mini -F language=mr \
-F diarize=true -F speakers=2Response
{
"text": "Hello sir, mi Pooja bolte Samarth Sky project madhun. Ha bola, kay aahe sanga...",
"segments": [
{ "start": 0.0, "end": 4.2, "speaker": 1 },
{ "start": 4.5, "end": 7.1, "speaker": 2 }
],
"x_minicrow": { "speakers": 2 }
}- The transcript is unchanged — labels are not interleaved into
text; join them on the times. - If the speaker pass fails you still get the transcript, and the ₹3.50 is not charged.
- A two-channel WAV with one party per channel is transcribed channel by channel, billed once by the length of the call — and
diarizeis then skipped and not charged.
Lark Live
Live speech to text over a WebSocket
Preview, enabled per account. On any other account the upgrade answers 503 tier_not_deployed. Ask us to enable yours.
Stream a call's audio as it happens. Every turn is answered twice: a fast draft the moment the turn ends, then a checked final. Everything is decided in the query, before the socket opens.
wss://api.minicrow.com/v1/audio/transcriptions/live?model=lark-mini&language=mr&script=latin&encoding=mulaw&sample_rate=8000
Authorization: Bearer mc_YOUR_KEY| Parameter | Values |
|---|---|
| language | Required. Any of the 21 languages |
| script | latin or native — required for every Indian language, Hindi included |
| encoding | linear16 (default) · mulaw · alaw |
| sample_rate | 8000 · 16000 (mulaw and alaw are 8000 only) |
| endpointing_ms | 300–1000, default 400 — how much silence ends a turn |
| vad | server (default) · client — you send {"type":"turn_end"} |
What you receive
{"type":"session.started","session_id":"6f1c…","model":"lark-mini","lane":"indic","language":"mr","script":"latin"}
{"type":"speech.started","turn":0,"start":0.42}
{"type":"transcript.partial","turn":0,"stage":"draft","text":"हो सर फ्लॅट चा बुकिंग अमाउंट पन्नास हजार आहे","start":0.42,"end":3.18}
{"type":"transcript.final","turn":0,"text":"Ho sir, flat cha booking amount 50 hazar aahe.","start":0.42,"end":3.18,"script_honoured":true}
{"type":"session.ended","audio_seconds":124.37,"turns":31,"billed_seconds":125,"cost":69.4445,"cost_currency":"INR_paise"}Send audio as binary frames (20 ms to 1 s each) and {"type":"end"} when done — do not just hang up, or the last finals are lost. ₹20 / hour, billed per second received, rounded up once per session.
Pica
Text to speech (POST /v1/audio/speech)
Returns audio/wav. The charge rides in the response headers, so you can read what a synthesis cost without parsing the audio. Billed per character of input, at most 5,000 characters per request.
curl https://api.minicrow.com/v1/audio/speech \
-H "Authorization: Bearer mc_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{ "input": "नमस्ते, आज मौसम बहुत अच्छा है।", "voice": "david", "mode": "natural" }' \
--output hello.wavResponse headers
X-Voice: david X-Mode: natural X-Lane: standard
X-Cost-Paise: 10.8000 X-Cost-Currency: INR_paise
X-Duration-S: 2.05 X-Sample-Rate: 24000| Field | Notes |
|---|---|
| voice | A built-in voice or one of yours (mcv_…). Omitted: david for Devanagari text, robert otherwise |
| mode | normal or natural (default) — delivery style |
| lane | standard (default, Hindi and English) or expressive (Hindi, with emotions) |
| emotion | expressive lane only: neutral · happy · sad · angry · disgust · fear · surprise |
| stream | true streams length-prefixed WAV frames, one per sentence |
| first_clause | true speaks the first clause on its own — first audio about 437 ms sooner at p50 |
| response_format | On a stream, "mulaw_8k" sends 8 kHz G.711 μ-law frames for a phone line |
Streaming to a phone line
{
"input": "आज शाम चार बजे, राजेश के साथ आपकी कॉल तय है।",
"voice": "david",
"stream": true,
"first_clause": true,
"response_format": "mulaw_8k"
}emotion on the standard lane is a 400, not an ignored field, and an English voice will not read Devanagari (400 script_mismatch) — a flat or nonsense reading billed as if it worked is worse than a refusal.
Voices
List voices (GET /v1/audio/voices)
Eleven built-in voices — seven Hindi, four English — each pinned to a fixed reference clip, so the voice you shipped last month is the voice you get today; plus any voices your account has made.
{
"object": "list", "built_in_count": 11, "own_count": 1,
"data": [
{ "id": "david", "name": "David", "language": "hi", "gender": "male", "built_in": true },
{ "id": "mcv_k4t2…", "name": "Anil", "language": "hi", "kind": "generated", "built_in": false }
],
"modes": ["normal", "natural"],
"lanes": ["standard", "expressive"]
}Osprey Live
Real-time voice agent over one WebSocket
Live — Hindi and English, on every account. Marathi, Tamil, Telugu and more are coming soon.
You stream the caller's audio in. MiniCrow hears each turn, decides what to say and which of your tools to call, and speaks the reply back down the same socket while it is still being written. About ₹66 per call-hour for hearing, brain and voice together, as an illustration.
wss://api.minicrow.com/v1/agent/live?model=osprey-live&language=hi&encoding=mulaw&sample_rate=8000&output=mulaw_8k&voice=david
Authorization: Bearer mc_YOUR_KEY- The upgrade succeeds and you receive
session.started. - Within 10 seconds, before any audio, send
session.configurewith the agent spec. - On
session.configured, start sending the caller's audio as binary frames — the caller's track only, on one channel.
session.configure
{
"type": "session.configure",
"agent": {
"spec_version": "osprey-live-prompt/0.1",
"persona": { "name": "Asha" },
"languages": { "caller_speaks": ["hi", "en"], "agent_speaks": "hi", "agent_script": "latin" },
"brief": "You are Asha, the appointment desk voice for Leafview Family Clinic. Speak one short sentence of romanised Hindi per reply. To check or book an appointment, call the tool.",
"tools": [{ "name": "check_appointment_slots", "kind": "booking" }]
}
}One turn with a tool
{"type":"input.speech_started","turn":7,"start":41.232}
{"type":"response.started","response_id":"r_12","turn":7,"kind":"reply"}
{"type":"response.audio.started","response_id":"r_12","clause":0,"kind":"speech","text":"Ji, time dekh leti hoon."}
{"type":"tool.call","response_id":"r_12","call_id":"call_7f3aQ2mX9kLp","name":"check_appointment_slots","arguments":{"doctor":"Dr. Arjun Rao","date":"2026-03-04"}}
{"type":"response.done","response_id":"r_12","turn":7,"status":"completed","passes":2,"tool_calls":1}
{"type":"turn.usage","turn":7,"brain":{"cost":3.3},"voice":{"characters":93,"cost":8.37},"cost":11.67,"cost_currency":"INR_paise"}Answer a tool.call
{ "type": "tool.result", "call_id": "call_7f3aQ2mX9kLp", "output": { "status": "sent" }, "is_error": false }Binary frames between response.audio.started and response.audio.done are the agent's voice, paced for playback. Send response.say to speak fixed text, response.cancel to stop, and end to finish. Every field is in the full reference.
Lark-V
Video summaries (POST /v1/video/summaries)
One upload returns a summary and a seekable, timestamped timeline. lark-v-large watches the video natively — it sees the frames and hears the speech. At most 20 MB.
curl https://api.minicrow.com/v1/video/summaries \
-H "Authorization: Bearer mc_YOUR_KEY" \
-F file=@clip.mp4 -F model=lark-v-large -F effort=highResponse
{
"model": "lark-v-large",
"summary": "A colorful test pattern is shown while a voiceover announces a meeting in Pune.",
"timeline": [{ "t": 0, "text": "A test pattern displays while a voice says, \"Kal shaam 5:00 baje…\"" }],
"usage": { "prompt_tokens": 288, "completion_tokens": 183, "cost": 9.5278, "cost_currency": "INR_paise" }
}| effort | What you get |
|---|---|
| mid | The key moments |
| high (default) | A moment every few seconds |
| max | Every distinct moment |
Retrieval
Embeddings and rerank
POST /v1/embeddings returns dense and sparse vectors from one call — the sparse half is what hybrid search needs. At most 256 inputs per request. POST /v1/rerank scores a shortlist against a query and is billed on query plus documents.
/v1/embeddings
{ "model": "minicrow-embed", "input": ["kal meeting hai Pune me", "tomorrow there is a meeting"] }/v1/rerank
{
"model": "minicrow-rerank",
"query": "Pune meeting kab hai",
"documents": ["Delhi ka flight subah 6 baje", "Pune me meeting kal 3 baje hai"],
"top_n": 2
}Response shape
The cost of every call, in paise
usage.cost is always MiniCrow's charge, decimal, in INR_paise — the same number in a stream as without one. A short call costs a fraction of a paisa, which is why the field is not rounded.
"usage": {
"prompt_tokens": 88,
"completion_tokens": 60,
"cost": 0.6458,
"cost_currency": "INR_paise"
}| Endpoint | Where the cost is |
|---|---|
| Chat, transcription, video, embeddings | usage.cost in the JSON body |
| Text to speech | X-Cost-Paise response header |
| Live transcription | cost on session.ended |
| Voice agent | cost on every turn.usage |
cost_known: false means no price was reported and the lane has no flat rate — nothing was charged. It is a gap we show you, not a discount.
Errors
HTTP semantics and what you should do
Every failure is the same JSON envelope — {"error":{"message":"…","type":"…","code":"…"}} — so a client can always branch on error.code.
| HTTP | code | When | What to do |
|---|---|---|---|
| 200 | — | Success | Read usage.cost |
| 400 | unknown_model / unknown_mode | No such tier, or a mode that tier does not offer | Fix the model id; the message lists valid ones |
| 400 | context_too_long | Prompt plus max_tokens exceeds the lane | Trim the input or lower max_tokens |
| 400 | unsupported_audio / empty_audio | A container we cannot measure, or zero seconds | Send WAV, OGG/Opus, MP3, M4A or WebM |
| 401 | invalid_api_key / key_revoked | Missing, wrong or revoked key | Stop and alert an operator; do not retry |
| 402 | insufficient_credit | Balance at or below zero | Top up, then retry |
| 402 | key_limit_reached | This key hit its own limit | Raise the cap or use another key |
| 403 | account_suspended | The account is suspended | Contact support |
| 413 | file_too_large | Audio over 25 MB, video over 20 MB | Split the file |
| 429 | rate_limited | The lane is busy, or too many requests | Back off and retry, or use another mode |
| 503 | model_unavailable / model_timeout | The model could not answer in time | Retry after Retry-After; not charged |
| 503 | tier_not_deployed | The tier or mode is not serving | Use the tier the message names |
A model provider's own 401, 402 or 404 is answered as model_unavailable, never as your credential or credit problem. The full reference lists every code.
Rate limits
Limits, retries and 503s
50 requests per second sustained, burst 100, per client IP. That is far above what a real integration sends; it exists to bound the cost of refusing bad keys. A busy lane answers 429 rate_limited — back off exponentially (1 s → 2 s → 4 s, capped around 30 s) and retry.
MiniCrow never answers 502 or 504: a gateway failure comes back as 503 with the same error.code, a Retry-After: 5 header, and X-MiniCrow-Origin-Status naming what the gateway decided. model_unavailable is recorded but not charged.
For AI coding agents
Integrate MiniCrow using Cursor, Antigravity or your own agent
We publish a single, opinionated agent.md that lays out everything a coding agent needs to integrate MiniCrow end to end: which model for which job, the request shapes, streaming, and the language and script rules that decide accuracy. Drop it at the root of your project — modern agents auto-load AGENTS.md from there.
Auto-detected by
Two-second install (run in your project root)
curl -fsSL https://www.minicrow.com/agent.md -o AGENTS.md
git add AGENTS.md && git commit -m "docs: add MiniCrow AGENTS.md"Then prompt your agent: “Integrate MiniCrow into this app — read AGENTS.md, store the API key as MINICROW_API_KEY, and add a transcription endpoint using lark-mini.” Agents follow the file end to end without further context.