# MiniCrow for AI agents

This file is for the agent (or engineer) integrating MiniCrow into an existing system. It says what each model is, when
to use it, the exact request shape, the rules that decide accuracy, and what to avoid. Everything here is measured on
MiniCrow's own benchmarks; where a number is directional it says so. The full reference is `/docs` (the API reference);
this page is the short, prescriptive version. Base URL `https://api.minicrow.com`, header `Authorization: Bearer <key>`,
OpenAI-compatible wire format. Every response carries the price of the call in paise.

## The models, in one table

| Model | What it does | Call it when | Do not call it when |
|---|---|---|---|
| `osprey-flash` (`:speed` · `:intelligence` · `:max` · `:auto`) | Text chat and tool calling, OpenAI-compatible | You need an LLM answer, a JSON extraction, a tool call from text | You have audio (use Lark) or a live phone call (use Osprey Live) |
| `osprey-pro` (`:speed` · `:intelligence` · `:max` · `:auto`) | The same, larger model | Long reasoning, hard extraction, code | Simple replies — `osprey-flash` is cheaper and faster |
| `osprey-flash-lite` — **coming soon** | The smallest Osprey | Not yet: every call is refused `503 tier_not_deployed`. Use `osprey-flash:speed` | Until `GET /v1/models` lists it `available: true` |
| `lark-nano` · `lark-mini` · `lark-large` | Speech to text from a recording, 21 languages, code-mixed. Nano is coming soon, mini the default, large the most accurate | You have a finished recording (call, voice note, meeting) | You need words while the caller is still speaking (Lark Live) |
| `lark-mini` live (`/v1/audio/transcriptions/live`) | Speech to text over a WebSocket, turn by turn, ~0.7 s after the caller stops | A live call where your own system decides what to say | You want MiniCrow to answer the caller (Osprey Live) |
| `osprey-live` (`/v1/agent/live`) | A complete voice agent over one WebSocket: hearing → your instructions and tools → speech. It also sees: send pictures from the caller's camera beside the audio | A phone agent that books, checks, sends, hands over; an assistant on glasses, a robot or a kiosk that answers about what its camera shows | Batch work, or languages other than Hindi and English (coming) |
| `pica-nano` · `pica-small` · `pica-large` | Text to speech, Hindi and English, streaming, telephone format; large is the most natural | You already have the words and need audio | — |
| `lark-v-nano` · `lark-v-mini` · `lark-v-large` | Video → summary and timeline JSON (what is said and seen, by second). Large is the most accurate: it sees and hears together; nano is coming soon | A recording with a picture that matters | Audio only (Lark is cheaper) |
| `minicrow-embed` (`/v1/embeddings`) · `minicrow-rerank` (`/v1/rerank`) | Embeddings (dense and sparse) and reranking, for search and RAG | You index or search your own text | — |

Every generative model above answers as MiniCrow's: `osprey-flash`, `osprey-pro`, or your Osprey Live agent's persona.

## Rules that hold everywhere

1. **Declare the language.** `language` on every speech call (`hi`, `mr`, `en`, …). MiniCrow never guesses; the
   declaration picks the lane and tunes the transcription to that language. A wrong or missing language costs more
   accuracy than anything else.
2. **Declare the alphabet for Indian languages.** `script=latin` (romanised, as people type in chat) or `script=native`.
   The transcript follows it; English words inside a sentence stay in English either way.
3. **Telephony audio is fine as it is.** 8 kHz μ-law/A-law or 16-bit PCM straight from the trunk. Do not upsample or
   "enhance" it; MiniCrow's speech lanes were measured on 8 kHz calls.
4. **One speaker per channel.** For calls, send each party on its own channel or its own session. Mixed audio makes the
   agent hear itself, and no model separates two people in one channel well.
5. **Hints are optional and small.** Accuracy is built to hold with none. Send the names this recording or call will
   contain (`vocabulary`), not a glossary of everything you sell; a long list costs more on every turn and can add words
   that were never said.
6. **Read the response.** `x_minicrow` says which lane served, why, and the exact price. Log it; it is your bill and your
   audit trail.

## Lark (recorded speech → text): best practice

Endpoint `POST /v1/audio/transcriptions`, multipart: `file`, `model=lark-mini`, `language`, `script`, optional
`domain`, `vocabulary`, `abbreviations`, `prompt`, `diarize`.

- Send WAV/OGG-Opus/MP3/M4A/WebM. The bill is per second, read from the file header.
- Two-channel calls: `diarize=true` with a stereo file labels the parties from the channels — the cheapest and most
  accurate speaker labelling MiniCrow offers. One channel: `diarize=true` runs the speaker pass (slower, ₹ per hour).
- The hint fields, in order of value: `vocabulary` (names spelled as you want them), `abbreviations` (`SHORT = long`),
  `domain` (one line, who talks to whom), `prompt` (anything else, background only). Measured: eight names and a
  one-line domain moved accuracy within the margin on read speech and helped only where the names were actually spoken.
- Never paste your agent script, FAQ or a previous transcript into `prompt`: on a near-silent clip Lark can return
  it as the transcript. MiniCrow refuses a transcript that is a copy of one of Lark's own reference lines; it cannot
  know yours.
- Long recordings: `lark-mini` has no duration cap below 25 MB. `lark-nano` is coming soon.
- A WAV that certainly holds no speech — silence, a steady tone or hum, line noise, a busy or ringing tone — comes back
  `text: ""` with `x_minicrow.no_speech: true`, not billed. On a call split by channel, a channel with no speech
  comes back as an empty segment with `no_speech: true`. A tone that switches on and off can still be transcribed.
- Two channels: channel 1 is the left one. A split call returns one segment per channel, each with its own
  `script_honoured` and `script_repaired`; the call is honoured only when every channel is.
- A transcript that loops is never returned: it is transcribed again without your hints, or the loop is kept once;
  `x_minicrow.repetition` says which (`retried_without_hints` or `collapsed`).
- `verify_vocabulary=true` checks every vocabulary name the transcript writes against a second transcription made
  without the vocabulary, and replaces a name that was not heard with what was. It costs one more transcription of the
  recording, only when a name was written; `x_minicrow.vocabulary_check` lists what was confirmed, removed or left.

## Lark Live (streaming speech → text): best practice

Endpoint `GET /v1/audio/transcriptions/live` (WebSocket). Query: `model=lark-mini`, `language`, `script` (required
for Indian languages), `encoding` (`linear16`|`mulaw`|`alaw`), `sample_rate` (8000|16000), `endpointing_ms`
(300–1000), `vad` (`server`|`client`), `domain`, `vocabulary`, `context` (repeatable).

- Send audio at real-time pace in 20 ms–1 s frames. Read `transcript.partial` (stage `draft`, ~0.7 s after the caller
  stops) to act fast and `transcript.final` (~1.5 s later) to store; the final corrects the draft.
- `endpointing_ms` is the trade: 300 answers sooner and splits sentences; 600–800 keeps a slow speaker's sentence whole.
  Default 400. Use `vad=client` only if your telephony stack already detects turns.
- `vocabulary` (up to 40 terms) is used on every turn AND is boosted inside the draft for `hi`, `mr` and
  `en` — the languages the boost is measured on; the other 18 use it as a spelling hint only, without the boost, until
  measured. List names, never everyday words; for `hi`/`mr` add the Devanagari spelling as its own term (`Orbit, ऑर्बिट`). Measured:
  Hindi/Marathi boosted names caught 5 of 7 against 2 of 7 without, word error unchanged; English 13 of 13 against 12,
  word error 14.3 → 12.4.
- `hi` and `mr` drafts are locked to Devanagari (English words in English letters stay), vocabulary or not.
- Reconnect mid-call by opening a new session with the last finals in `context`.
- Speaker labels are not offered live; one session per channel is the way.

## Osprey Live (voice agent): best practice

Endpoint `GET /v1/agent/live` (WebSocket). Query: `language` (`hi`|`en`), `voice`, `encoding`, `sample_rate`,
`output` (`mulaw_8k` for a phone line), `endpointing_ms`, `barge_in`, `transcripts`. Then one `session.configure`
message with your agent spec, then audio in, audio + events out. You run every tool; MiniCrow sends `tool.call` and waits
for your `tool.result`.

The spec decides whether tools get called. Measured on 99 turns where a tool was required (invented clinic, car service,
courier and property calls, Hindi and English, plus real Marathi calls): the default brain made 59 % of them; with the
caller's words heard perfectly the same brain made 81 %, and the larger brain 89 %. Hearing, then spec wording, are the
levers you own.

**The spec, field by field, and the rule for each**

| Field | Rule |
|---|---|
| `brief` | Who the agent is, how it speaks (language, script, one short sentence, numbers as digits), the facts it may state, the goal of the call in order. Say "call the tool for anything not stated here". Keep it short: it is sent every turn |
| `tools[].when_to_call` | Start with "CALL THIS", name situations not test sentences, say BEFORE ("CALL THIS BEFORE saying any time"), close the escape ("never answer from memory") |
| `tools[].when_not_to_call` | For actions: "already done on this call", "number still being dictated" |
| `tools[].parameters` | Require only what the tool cannot run without; ask optional details after the result. A `lookup` or `booking` with more than 3 required fields waits for all of them. Describe a phone field as a value ("caller_id when they say to use this number, otherwise the number they give"), never as steps ("dictated and confirmed") |
| `tools[].kind` | `lookup` (reads), `booking` (checks, commits nothing), `action_needs_confirmation` (the caller will notice), `other` |
| `tools[].confirm: "runtime"` | Put it on every action whose arguments the caller dictates (phone, name, email). The first call is read back digit by digit and a misheard value is asked again; the call reaches you only after a yes |
| `tools[].acknowledgement` | 3–5 spoken words that promise nothing ("Ji, time dekh leti hoon.") |
| `examples` | One invented example per tool you most need called, one where the lookup does not answer, one number dictation. Everything in them invented; never a real number |
| `hearing.vocabulary` | Names your callers say: people, products, places, codes, up to 100, with `spoken` Devanagari forms for Hindi and Marathi callers (`{"term":"Orbit","spoken":["ऑर्बिट"]}`). Names, never everyday words. Tool enum values are boosted on their own. Boosted in the hearing for `hi`, `mr`, `en`; Hindi and Marathi hearing is locked to Devanagari |
| `speech.reprompt`, `greeting_words`, `never_claim_before_result`, `block_before_result` | What to say on an unreadable turn; words that mean a greeting; outcome words the agent may not say before a tool result |
| `history` | The greeting you already played, or the call so far on a reconnect |

`session.configured.warnings` checks your spec against these rules and names the field. Fix them before going live.

**The camera (vision).** Send a picture from the caller's camera every 1–3 seconds as
`{"type":"input.image","data":"<base64 JPEG, PNG or WebP>"}`, after `session.configured` (which carries `vision` on a
session that takes pictures). At most 192 KB a picture — 640 to 1024 pixels wide is plenty; at most one a second is
kept. When the caller finishes a turn, the reply is written with up to 3 of the pictures taken while they spoke, so
"what is this?" is answered about what the camera showed. Only the turn's own pictures are seen; what the agent said
about them stays in the history as text. Pictures are billed as brain input tokens — about 1,100 each.

**Runtime you can rely on** (whatever the prompt says): no outcome claim before a tool result; no second greeting; an
identical successful action is not repeated; at most 3 brain passes and 2 tool calls per caller turn; asked what it is or
who made it, it answers only as your persona; a tool call the agent writes into its words instead of making it is never spoken; runtime confirmation as above — and a yes to the read-back releases the held call on the next
reply; a phone number the caller never said — a placeholder such as
9876543210, or one copied from your examples — is never sent to a tool: the agent is told to use `caller_id` (the caller
said "use this number") or to ask for the number.

**Latency shape:** first audio ≈ 2.3–2.6 s after the caller stops on a cache read; the first 3–4 turns of a call are
warmed at connect. A tool turn speaks the acknowledgement, calls the tool, and answers from the result. Play audio as it
arrives; do not buffer.

**Choosing the brain:** the default brain is the cheaper one (our measured 59 % → 66 % with hearing vocabulary on the
turns above); the larger brain scores 74–76 % at about 1.5× the brain price. Ask MiniCrow to switch your account if
required tool calls matter more than the difference.

## Osprey (text chat): best practice

`POST /v1/chat/completions`, `model` = `osprey-flash` | `osprey-flash:speed` | `osprey-flash:intelligence` |
`osprey-flash:max` | `osprey-flash:auto` | `osprey-pro` (same modes). `osprey-flash-lite` is coming soon.

- `:auto` chooses the lane per request from the request itself (tools, length, script) and may step up once after a
  verifiable failure; both attempts are charged. Name a mode when you need a fixed price.
- Send `X-MiniCrow-Session` with a stable id per conversation so MiniCrow keeps the same lane across turns.
- Tools: OpenAI `tools` + `tool_choice`. Keep schemas tight (`enum` where the values are known); the response reports
  `tool_calls_valid`.
- Streaming: `stream: true`; the final chunk carries `usage` and the cost.
- Thinking counts toward `max_tokens` and is billed; if it uses the whole budget the answer is empty with
  `finish_reason: "length"`. Give thinking modes 2,000+ tokens. On `:speed`, `reasoning_effort: "none"` switches
  thinking off for that request: about 3 s instead of 8 and never empty — good for summaries and extraction, worse for
  calculations. `:intelligence` always thinks.
- The model presents itself as `osprey-…` by MiniCrow.

## Embeddings and reranking (search and RAG): best practice

- `POST /v1/embeddings` with `model=minicrow-embed` and `input` (a string or up to 256 strings): each text gets a 1,024-number
  dense vector and sparse lexical weights in the same response. Keep both — hybrid search (dense + sparse) finds the
  exact names and codes that dense vectors alone miss. Send text, never token-id arrays.
- `POST /v1/rerank` with `model=minicrow-rerank`, `query`, `documents`, `top_n` (Cohere's shape): rerank your top
  20–50 candidates from search, not the whole collection. Billed on the query and the documents.

## Pica (text → speech): best practice

`POST /v1/audio/speech` with `model`, `voice`, `input`, `response_format` (`mulaw_8k` for a phone line), `stream: true`
for first audio as soon as it is ready.

- Pick the tier: `pica-nano` (₹21/h) for the eleven built-in voices, your own cloned voices, emotions and
  `first_clause`/`chunk`; `pica-small` (₹32/h) for its 25 Hindi and English voices; `pica-large` (₹60/h) when the voice
  has to sound the most natural. List a tier's voices with `GET /v1/audio/voices?model=…`.
- A built-in or cloned voice is always spoken by `pica-nano`, whatever `model` says; `X-Model` tells you which tier
  answered. `lane`, `emotion`, `chunk` and `first_clause` on small/large are a 400, not ignored.
- `429 tier_busy` on small/large: retry shortly, or fall back to `pica-nano`.
- Write the text as it should be said: digits for numbers, Devanagari for a Hindi voice. The bill is per second of
  audio received, so shorter wording costs less.
- Voice cloning (`POST /v1/audio/voices`) needs the speaker's consent; the reference explains what is stored.

## Lark-V (video → timeline): best practice

`POST /v1/video/summaries` with the file, `model` (`lark-v-mini` | `lark-v-large`; `lark-v-nano` is coming soon), `language`,
`script`, `summary_language`, `effort` (`mid`|`high`|`max`). The answer is JSON: a summary and a timeline of
`{t, text}` in seconds. `lark-v-large` is the most accurate — it sees the picture and hears the sound together. Use
`effort=mid` unless frames matter. Billed on what the clip used; every response carries its cost.

## Integration checklist

- [ ] Base URL and key set; `x_minicrow` logged from every response.
- [ ] `language` (and `script` for Indian languages) declared on every speech call.
- [ ] Telephony audio sent as-is (8 kHz, one party per channel).
- [ ] Live: partials used to act, finals stored; `endpointing_ms` tuned on your own calls.
- [ ] Osprey Live spec: `session.configured.warnings` empty; `confirm: "runtime"` on dictated-value actions;
      `hearing.vocabulary` filled with names (Devanagari `spoken` for hi/mr); examples invented.
- [ ] Tools answered within `tool_timeout_ms`; results short JSON; `call_id` echoed.
- [ ] Reconnect paths implemented (close 1012/1013 → new session with `history` / `context`).
- [ ] Osprey Live with a camera: pictures sent after `session.configured`, one every 1–3 s, each under 192 KB.

## Prices (customer, list; ₹)

- **Lark**, per hour of audio: nano ₹10, mini ₹20 (batch and live), large ₹30 — ₹45 when an Indian language is
  declared, ₹60 for `lark-large:max`.
- **Osprey Live**: hearing ₹20 an hour + brain per token by type (≈ ₹13.5 per call-hour on the default brain) + voice
  ₹9 per 10,000 characters (≈ ₹32 per call-hour) ≈ ₹66 per call-hour all-in. Camera pictures add brain input tokens:
  about 9 paise on a turn with three.
- **Pica**, per hour of audio: nano ₹21, small ₹32, large ₹60.
- **Osprey chat, embeddings and rerank**: per million tokens — `GET /v1/models` lists every rate.
- **Lark-V**: billed on what each clip used.

`GET /v1/models` and the price in every response are the authority.
