Build Without
Limits
Access the best AI models, run them reliably, and scale globally — all through one powerful API.

Models
take you further.
- One APIText, speech, video, retrieval
- Routing you can readEvery call says why
- Priced per callCost in the response
- PrepaidZero balance, hard stop
Language models
Three tiers. Pick the one that fits the job.
What you buy is a lane, and the lane is what stays stable while the model behind it is repriced, re-quantised or replaced. Every response still tells you which lane answered, the rule that picked it, and what the call cost — a router that hides its choice cannot be debugged.
Flash Lite
osprey-flash-lite
The cheap one that still sees and hears.
One model, two speeds: thinking off for volume, thinking on when the answer has to be right. It takes text, images, audio and files, which for its price is the surprise of the catalogue — but it will not take tools.
Good for
- High-volume classification
- Tagging and routing
- Cheap extraction passes
- Describing an attachment
Modes
:speed₹9.5 · ₹35.9Thinking off. For volume.
Takes text · images · audio · files
:intelligencedefault₹9.5 · ₹35.9Thinking on. The default.
Takes text · images · audio · files
₹ per million tokens, in · out
Priced and in the catalogue. The self-hosted endpoint behind it is not deployed yet.
Flash
osprey-flash
The workhorse. Nearly everything belongs here.
Three lanes across a 3× price range, with tools on every one of them. This is the tier the router was designed around: fast and cheap for short exchanges, the middle lane the moment tools appear, and a premium rung you can name when an answer has to hold up.
Good for
- Chat and assistants
- Agent loops with tools
- Structured extraction
- Reading screenshots
- Summarising a long document
Modes
:speed₹6.9 · ₹19.0Short exchanges, low latency.
Takes text only
:intelligencedefault₹7.9 · ₹26.4The default. Tools, images, most work.
Takes text · images
:max₹21.1 · ₹126.7When it has to hold up.
Takes text · images · files
₹ per million tokens, in · out
Pro
osprey-pro
A different price class, not a nicer Flash.
Reach for Pro when the work is a long agent loop with real consequences, or when a flash-class model has already been tried and failed. Its middle lane costs roughly nineteen times Flash's, which is the honest way to describe it: this is not an upgrade, it is a different budget.
Good for
- Multi-turn agent loops
- Long-context reasoning
- Code and analysis that must be right
- Audio and video in one prompt
Modes
:speed₹79.2 · ₹396.0The widest input of any lane.
Takes text · images · audio · files
:intelligencedefault₹147.8 · ₹464.6The default. Deep reasoning.
Takes text only
:max₹211.2 · ₹1,056.0The top of the catalogue.
Takes text · images · files
₹ per million tokens, in · out
Every language lane, side by side
1M+ context, every lane| Lane | Best for | Takes | Tools | ₹/Mtok in | ₹/Mtok out |
|---|---|---|---|---|---|
| osprey-flash-lite | Volume classification | text · image · audio · file | 9.5 | 35.9 | |
| osprey-flash:speed | Short chat, low latency | text | 6.9 | 19.0 | |
| osprey-flash:intelligence | Tools, images, most work | text · image | 7.9 | 26.4 | |
| osprey-flash:max | Answers that must hold up | text · image · file | 21.1 | 126.7 | |
| osprey-pro:speed | Widest input, fast | text · image · audio · file | 79.2 | 396.0 | |
| osprey-pro:intelligence | Deep reasoning | text | 147.8 | 464.6 | |
| osprey-pro:max | Long agent loops | text · image · file | 211.2 | 1,056.0 |
Flash-class models collapse on multi-turn tool use.
On BFCL v4's multi-turn split, flash-class models score 13–36% where pro-class models score 60–68%. It is not a uniform five-point tax — it is a cliff, and it is why a wrong cheap route breaks an agent loop while a wrong expensive one only costs markup.
Berkeley Function-Calling Leaderboard v4 — published third-party, not our measurement
We do not publish a benchmark board of our own, because we do not have one yet and borrowing someone else's would be worse than saying so. No public board puts these lanes side by side, and not one of them reports a Hinglish or Indic figure — which is most of the traffic this gateway was built for. Every call's route reason and cost is logged, and that is what our own numbers will be built from.
Speech, video and retrieval
The rest of the catalogue, built on Indian audio.
Transcription, video summaries, speech and vector search — the same key, the same prepaid balance, the same cost in the response.
Lark
Speech to textIndian-language transcription, code-mix included. Measured on real 8 kHz call audio, not on studio recordings.
Good for
- Phone-call transcription
- Voice notes and meetings
- Hinglish and Indic speech
- Support-call QA
- Anything recorded at 8 kHz
POST /v1/audio/transcriptions
- lark-nanoNot availableSelf-hosted₹7.07per hour of audio
- lark-miniDefault lane₹7.07per hour of audio
- lark-largeWider output budget₹42.50per hour of audio
Lark-V
VideoA summary and a seekable timeline from one upload. The large tier watches the clip natively — it sees the frames and hears the speech, in the script it was spoken in.
Good for
- Video summaries
- A seekable timeline of a clip
- Screen recordings and demos
- Ad and creative review
- Spoken content inside video
POST /v1/video/summaries
- lark-v-nanoNot availableTiled framesvision + speech+20% on the sum
- lark-v-miniNot availableTiled framesvision + speech+20% on the sum
- lark-v-largeNative video, sees and hearsupstream + 20%measured per clip
The large tier is live. The two frame-tiling tiers are not deployed and say so rather than upgrading you to it.
Pica
Text to speechEleven ready voices — seven Hindi, four English — each a sha256-pinned reference. Two lanes: standard for both languages, expressive for Hindi with emotion. Two delivery modes on each.
Good for
- Voice notifications and IVR
- Hindi narration with emotion
- Audio versions of written content
- Product and demo voiceover
- Accessibility read-aloud
POST /v1/audio/speech
- pica-ministandard and expressive lanes₹36per 10,000 characters
- pica-largeNot availableAnnounced, not wired up yet—price not set
Embeddings and rerank
LiveDense and sparse vectors from one call, and a cross-encoder that reranks a shortlist. Unbranded on purpose — this one has not been given a name yet.
- Search over your own documents
- RAG retrieval
- Deduplication
- Reranking a shortlist
POST /v1/embeddings · POST /v1/rerank
₹2.1
per Mtok · dense + sparse
Prices are the catalogue's seed rates, at a 20% markup over what a call costs us and ₹88 to the dollar — a configured rate, not a live FX feed. The operator panel changes any of them, and a change never rewrites a call already made.
Auto mode
The router tells you what it picked, and why.
It is rules, not a model — structural signals read off the request body in under two milliseconds, with no English keyword lists, because a keyword list scores Hinglish, Gujarati and Tamil traffic identically to noise. Every decision comes back with the same reason id it was recorded under, so a route you disagree with is a string you can search for rather than a mood you have to argue with.
H6:open_tool_loopNever switch mode inside an open tool loop
Thought signatures and thinking blocks bind a continuation to the model that began it. All three upstreams 400 when replayed elsewhere.
S:toolsTools present, or three turns deep, takes the middle lane
A wrong cheap route breaks the agent loop. A wrong expensive one only costs markup.
S:stickyA thread keeps the mode it started on
A switch throws away the prompt cache, and cache-read is a fraction of input price on every lane.
S:long_inputEight thousand tokens of input is a summarisation job
One long user turn with no tools is a different shape of work from a conversation.
H2A lane that cannot take the modality is removed
An image on a text-only lane is not a worse answer, it is an error.
EOne rung up, once, only on a verifiable failure
A tool call that cannot be executed is evidence. A truncated answer is not — that is a continuation problem in the same mode.
{
"model": "osprey-flash",
"x_minicrow": {
"requested_mode": "auto",
"served_mode": "intelligence",
"route_reason": "S:tools",
"cost_known": true
},
"usage": {
"prompt_tokens": 88,
"completion_tokens": 60,
"cost": 0.6458,
"cost_currency": "INR_paise"
}
}cost is in paise, to four decimals. A short call costs a fraction of a paisa — rounding each one to a whole paisa would report a busy month as free.
Never inside an open tool loop
Thought signatures and thinking blocks bind a continuation to the model that began it. Switching mid-loop is a 400, not a worse answer.
max is never chosen for you
The most expensive lane is reachable by naming it, or by one escalation after a verifiable failure. Never by the router deciding you meant it.
An explicit mode is obeyed exactly
Name a mode and it is never escalated. Serving something dearer than you asked for, and billing you for it, is an override, not a correction.
Lark · speech to text
The same accuracy, at a fifth of the price.
Built and measured on the audio Indian products actually have: 8 kHz telephone calls, code-mixed, mostly not in English. Every number here was taken on that, not on a studio benchmark.
- Phone-call transcription
- Voice notes and meetings
- Hinglish and Indic speech
- Support-call QA
- Anything recorded at 8 kHz
Compared against Sarvam
On 102 real 8 kHz Marathi calls the two scored 82.3 and 80.8 — level, inside the noise. The claim is the price, not the accuracy.
A fifth of the price, at the same accuracy
Lark mini sells at ₹7.07 per hour of audio. Sarvam's published beta price is ₹30.69. On 102 real 8 kHz Marathi calls the two scored 82.3 and 80.8 — level, inside the noise. You are not trading accuracy for the price.
Measured · 102 real calls
Telephone audio costs nothing
8 kHz call recordings were measured against wideband twice, and band-limiting cost zero accuracy both times. Lark takes 8 kHz to 48 kHz and the narrow end is not the cheap end.
Measured twice · 8 kHz vs wideband
Code-mix comes back in the script it was spoken in
Hinglish and romanised Indic are the traffic this was built on. Left alone these models write Devanagari for Hindi however it was said; Lark asks for it the way the speaker said it, so code-mixed speech returns in Latin script. Whichever you get, it is stated rather than discovered.
Live · stated default
A hint you can send, that cannot hijack the job
Send a name spelling or a domain term and it is appended to the instruction — never substituted for it. A prompt that could replace the task would turn transcription into general inference on an ASR-priced lane, which is somebody else's bill.
Live · appended, not substituted
Indian languages, first class
Marathi, Hindi and the rest are the evaluation set, not a footnote to an English benchmark. Every accuracy number quoted here was taken on Indian-language call audio.
Live today
Specified, not shipped. Not available today.
Speaker diarization with a timestamped timeline
Who spoke, when, as a timeline you can seek. Not built, and we would rather say so than ship a guess: the measured finding is that channel separation — one caller per channel at the recorder — is a far bigger lever on this audio than any diarizer, and that is where the work goes first.
Roadmap · not built
Transcribe and translate in one call
The models behind Lark can do it. Lark does not expose it yet, so it sits on this list as coming rather than in the list above as a feature.
Roadmap · not exposed yet
Pica · text to speech
Eleven voices, and seven emotions in Hindi.
Seven Hindi voices and four English, each frozen against a sha256-pinned reference so the voice you shipped last month is the voice you get today. Priced per character, with no per-seat tier.
- Voice notifications and IVR
- Hindi narration with emotion
- Audio versions of written content
- Product and demo voiceover
- Accessibility read-aloud
Compared against Sarvam and ElevenLabs
POST /v1/audio/speech
Seven, enumerated and enforced — an unknown one is a 400, not a silent fall back to flat delivery. Emotions live on the expressive lane, which is the Hindi one; the standard lane covers both languages and has none.
Seven emotions, on the expressive lane
neutral, happy, sad, angry, disgust, fear, surprise — enumerated and enforced. Send lane=expressive, the Hindi lane, and pick one. The standard lane has no emotions and says so with a 400 rather than returning a flat reading billed as though it had worked.
Live · lane=expressive
Eleven voices, live today
Seven Hindi and four English, each frozen against a sha256-pinned reference clip so the voice you shipped last month is the voice you get today. Two delivery modes on each, and an English voice handed Devanagari refuses rather than reading nonsense you would still be billed for.
Live · GET /v1/audio/voices
₹36 per 10,000 characters
Billed per character, the way every TTS vendor prices and therefore the way you will compare us. No per-seat tier, no minimum, and no separate charge for the voice. The cost of each synthesis comes back in the response headers, so you do not have to parse a WAV to find out what it cost.
Live · billed per character
Specified, not shipped. Not available today.
Bring your own voice
Generate a new voice from a written description, or clone one from a clean reference clip. Designed, priced and specified — including that cloning is a rights surface and what we do and do not verify — but not built.
Roadmap · designed, not built
Getting started
Change two things. Keep your client.
OpenAI-compatible because that is what every client library already speaks. The base URL and the key are the whole migration — the mode rides on the model id, so even that survives a library that has never heard of us.
from openai import OpenAI
client = OpenAI(
base_url="https://api.minicrow.com/v1", # ← 1
api_key="mc_96bf0550e045_…", # ← 2
)
r = client.chat.completions.create(
model="osprey-flash:auto",
messages=[{"role": "user", "content": "Aaj ka plan kya hai?"}],
)
print(r.choices[0].message.content)
print(r.model, r.usage.cost, "paise")Create a key
It is shown once. Only an argon2id hash is stored, so a lost key is replaced, not recovered.
Top it up
Keys are prepaid and hold a rupee balance. A key at zero gets a 402 before anything upstream is called.
Send a request
Every response carries the branded lane, the model that actually answered, the reason, and the charge.
Pricing
No plans. A balance, and a rate per model.
Keys are prepaid and hold rupees. Each call deducts what it cost, the charge comes back in the response, and a key at zero stops rather than surprising you with an invoice.
One markup, stated
Every model is priced as upstream cost plus 20%. A repriced upstream does not silently eat the margin, and a rate change never rewrites what was already charged — the rate is recorded on the call.
Charged per call, in paise
usage.cost is in the response, decimal, and it is our charge rather than an upstream figure. The same number in a stream as in a non-stream.
Prepaid, and it stops
A key holds a rupee balance. At or below zero it gets a 402 before the upstream is called, so a spent key never costs you anything. There is no postpaid billing.
A gap you can see
If an upstream reports no price and the lane has no flat rate, cost_known comes back false and nothing is charged. That is a gap we show you, not a discount we claim.
Seed rates
+20% markup| Model | Rate |
|---|---|
| osprey-flash-lite | 9.5 · 35.9₹ / Mtok in · out |
| osprey-flash : speed | 6.9 · 19.0₹ / Mtok in · out |
| osprey-flash : intelligence | 7.9 · 26.4₹ / Mtok in · out |
| osprey-flash : max | 21.1 · 126.7₹ / Mtok in · out |
| osprey-pro : speed | 79.2 · 396.0₹ / Mtok in · out |
| osprey-pro : intelligence | 147.8 · 464.6₹ / Mtok in · out |
| osprey-pro : max | 211.2 · 1,056.0₹ / Mtok in · out |
| bge-m3 | 2.1₹ / Mtok |
| lark-mini | 7.07₹ / hour of audio |
| lark-large | 42.50₹ / hour of audio |
| pica-mini | 36₹ / 10,000 chars |
Seed values at ₹88 to the dollar, which is a configured rate and not a live FX feed. The operator panel changes any row, and a change never rewrites a call already made.
Point your client at it and see what it costs.
Two lines of config, a prepaid key, and a response that tells you which lane answered, why it was chosen, and what you were charged for it.