₹21
per hour of audio
Affordable at scale
From ₹21 an hour on pica-nano; ₹32 on pica-small and ₹60 on pica-large, the most natural. Billed by the second.
Introducing
Natural, expressive, multilingual text to speech — with emotions, real-time streaming and voices of your own, for voice agents, companions and content.
Free in the Playground today · Public launch 1 October 2026
Voice AI is moving beyond robotic speech.Real voices. Real emotions. Real conversations.
For applications where how something is said matters almost as much as what is said.
Building an AI companion, a voice agent, an audiobook app or a game full of characters? The voice is the part you shouldn't have to build. Pica speaks your agent's words — with emotion, in Hindi, English and ten more Indian languages, streaming as it speaks — so your time goes into the harness: the product, the prompts, the tools and the memory.
You buildyour code
The harness
# one turn of a voice app
audio → Lark → text
text + prompt + tools → Osprey → reply
reply → Pica → voice
We runapi.minicrow.com
The senses
One keyOne billPrepaid in rupeesCost in every response
Press play: Aarav and Trisha, two of Pica's voices, and the same voice as a phone line hears it — made by Pica, in the language you pick.
Hear it in your language
Pica
Choose a scenario
Streaming · first clause in under a second
Pica's audio
Pica's own audio for the text above, in Aarav's voice.
आज शाम चार बजे, राजेश के साथ आपकी कॉल तय है।
Audio streams back sentence by sentence — each one plays while the next is being made.
A support agent that sounds sorry, a companion that laughs along, a story that turns frightening — on Pica's expressive lane every sentence can carry the feeling the moment needs.
On the expressive lane (Hindi, pica-nano): lane: "expressive" with an emotion.
₹21
per hour of audio
From ₹21 an hour on pica-nano; ₹32 on pica-small and ₹60 on pica-large, the most natural. Billed by the second.
7
emotions
Happy, sad, angry, fear, surprise, disgust or neutral — the same words, said the way the moment needs.
Live
sentence by sentence
The first sentence plays while the rest is still being made — no waiting for a whole file.
Yours
designed or cloned
Describe a voice and get one, or clone one from 10–20 seconds of speech you have the rights to.
12
languages
Hindi, English, Marathi, Tamil, Telugu, Gujarati, Bengali, Kannada, Malayalam, Punjabi, Odia and Assamese.
8 kHz
phone-line audio
μ-law frames for a telephone trunk, straight from the stream — for voice agents and IVR.
Pica against the three voice APIs developers compare it with: naturalness on MiniCrow's current benchmark, real-time streaming, and the published list price of an hour of speech.
Swipe the table to compare all four
| Benchmark | SarvamBulbul v3 | ElevenLabsFlash to v3 | GoogleChirp 3 HD voices | |
|---|---|---|---|---|
| Naturalness ↑1 | 92% | 88% | 92% | 91% |
| Real-time streaming | Yes | Yes | Yes | Yes |
| List price, per hour of speech2 | ₹21–₹60live price book | ≈₹151₹30 per 10,000 characters | ≈₹242–₹484$0.05–$0.10 per 1,000 characters | ≈₹145$30 per million characters |
MiniCrow's current benchmark measurements; results vary by language, voice, text, evaluation methodology and audio conditions. Sarvam, ElevenLabs and Google are third-party providers named for comparison only; MiniCrow is not affiliated with them.
POST /v1/audio/voices
The new voice's id is then just a voice on /v1/audio/speech.
curl
# design a voice from a description
curl https://api.minicrow.com/v1/audio/voices \
-H "Authorization: Bearer mc_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"name": "Anil", "language": "hi",
"description": "warm, unhurried, mid-forties"}'
# or clone one from 10–20 seconds of clean speech
curl https://api.minicrow.com/v1/audio/voices \
-H "Authorization: Bearer mc_YOUR_KEY" \
-F file=@reference.wav -F name=Anil -F language=hi \
-F attest=trueIllustrative scenes.
Lark handles the hearing. Osprey handles the reasoning and the tools. Pica speaks with expression. And Osprey Live brings the real-time experience together, on one phone socket.
A sentence, with feeling
POST /v1/audio/speech — the cost comes back in the headers.
Python
import requests
r = requests.post(
"https://api.minicrow.com/v1/audio/speech",
headers={"Authorization": "Bearer mc_YOUR_KEY"},
json={
"model": "pica-nano",
"input": "अरे वाह! आपका ऑर्डर आज ही पहुँच जाएगा।",
"lane": "expressive", # Hindi, with emotions
"emotion": "happy", # or sad, angry, fear, surprise …
},
)
open("reply.wav", "wb").write(r.content)
print(r.headers["X-Cost-Paise"], "paise")Stream sentence by sentence
stream: true sends each sentence as a playable WAV frame the moment it is made.
First words sooner
first_clause speaks the first clause on its own — 437 ms earlier at the median, measured.
Straight onto a phone line
response_format mulaw_8k streams 8 kHz G.711 μ-law frames a trunk takes as they are.
Billed by the second
The playing time of the audio you received — the cost is in the response headers.
Every field, limit and error is in the API reference.
Three tiers, one endpoint. Live prices per hour of audio, billed by the second, prepaid in rupees.

The expressive one.
Eleven pinned voices, seven emotions on Hindi, and the voices you design or clone.
₹21
per hour of audio · $0.219 an hour

More voices.
25 Hindi and English voices of its own — and Malayalam, Punjabi, Odia and Assamese.
₹32
per hour of audio · $0.333 an hour

Most natural
The most natural.
17 voices, each speaking Hindi and English — and Marathi, Tamil, Telugu, Gujarati, Bengali and Kannada.
₹60
per hour of audio · $0.624 an hour
You pay for the playing time of the audio you received, measured from the audio itself — ₹21 an hour on pica-nano is 0.5833 paise a second. A request that fails before any audio is not charged; a stream you hang up on is charged for the audio made up to that point. The same audio costs the same with or without streaming.
Voice interfaces should feel less like machines reading text, and more like actual conversations.
Public launch
1 October 2026.
Build › Deploy › Scale
Real voices. Real emotions. Real conversations. Try every Pica voice free in the Playground today — no account needed.

Pica is MiniCrow's text-to-speech (TTS) API: realistic, expressive and multilingual AI voices, with emotions, real-time streaming and voices of your own — for voice agents, AI companions, call centres, content, e-learning, games and enterprise voice applications.
Seven: happy, sad, angry, fear, surprise, disgust and neutral. Send lane: expressive — the Hindi lane on pica-nano — with an emotion; an emotion the lane doesn't know is refused with a 400, not read flat and billed anyway.
Yes. POST /v1/audio/voices designs a new voice from a written description, or clones one from 10–20 seconds of clean speech. Cloning requires attest=true — your statement that you hold the rights and have the speaker's permission — and every attempt is logged. Your voices belong to your account alone and speak on the standard lane.
Yes. With stream: true each sentence arrives as a complete, playable WAV frame as soon as it is made, and first_clause speaks the first clause of the first sentence on its own — 437 ms earlier at the median in our measurement. For phone lines, response_format mulaw_8k streams 8 kHz μ-law.
Hindi and English on every tier. pica-large also speaks Marathi, Tamil, Telugu, Gujarati, Bengali and Kannada; Malayalam, Punjabi, Odia and Assamese are spoken by pica-small. The language follows the script of the text you send.
₹21 an hour of audio on pica-nano, ₹32 on pica-small and ₹60 on pica-large ($0.219, $0.333 and $0.624) — billed per second of audio received, prepaid in rupees, no subscription. A request that fails before any audio is not charged.
Pica is the voice; you build the harness around it. Hear the caller with Lark, let Osprey decide and call your tools, and speak the reply with Pica, streaming — or use Osprey Live, which runs all three on one phone socket.
MiniCrow launches Pica publicly on 1 October 2026. You can try the voices free in the Playground today.
Pica speaks; the rest of MiniCrow thinks, hears and watches. Press play: hear a real call, a recording and its transcript — or watch Lark-V read a real clip.
Human-like conversational AI — fast, affordable LLMs for assistants, companions and agents, with tool calling.
Plan my Sunday — something relaxed.
Brunch at 11, a walk by the lake at 4. Shall I book the table?
₹6.9
per million input tokens · $0.072
Speech to text for India — Hindi, Marathi, Tamil and more, measured on real 8 kHz phone calls.
Hindi · 8 kHz phone call → text, in real time
₹20
per hour of audio · $0.208
Video understanding — a summary, a seekable timeline and the spoken words, from a single clip.
Per clip
billed on what each clip used
Real-time multimodal AI that sees, hears, thinks and talks — with tool calling, for AI calling, companions, robots, glasses and cars.
Caller: “Do you deliver on Sundays?”
Osprey Live: “Yes — until 9 pm. Shall I book you a slot?”
Semantic search and RAG — embed your documents, rerank a shortlist, and give Osprey your own knowledge to answer from.