Introducing

PicaRealistic AI voices that feel human.

Natural, expressive, multilingual text to speech — with emotions, real-time streaming and voices of your own, for voice agents, companions and content.

Free in the Playground today · Public launch 1 October 2026

Voice AI is moving beyond robotic speech.Real voices. Real emotions. Real conversations.

  1. Text
  2. Voice
  3. Emotion
  4. Conversation

For applications where how something is said matters almost as much as what is said.

You build the harness. We give it a voice.

Building an AI companion, a voice agent, an audiobook app or a game full of characters? The voice is the part you shouldn't have to build. Pica speaks your agent's words — with emotion, in Hindi, English and ten more Indian languages, streaming as it speaks — so your time goes into the harness: the product, the prompts, the tools and the memory.

You buildyour code

The harness

  • Your productThe app your users open — on the web, on a phone, on a device.
  • Prompts and personalityHow it thinks, what it says, how it sounds.
  • Tools and workflowsYour APIs and the actions it is allowed to take.
  • Memory and knowledgeWhat it remembers, and the sources it answers from.
  • EvalsHow you know it works, call after call.

# one turn of a voice app

audio → Lark → text

text + prompt + tools → Osprey → reply

reply → Pica → voice

Hear it. Real Pica voices, not a recording of a person.

Press play: Aarav and Trisha, two of Pica's voices, and the same voice as a phone line hears it — made by Pica, in the language you pick.

Open the Playground

Hear it in your language

Pica

Choose a scenario

from ₹21 / hour of audio3 tiers7 emotions on the Hindi lane

Streaming · first clause in under a second

Pica's audio

Pica's own audio for the text above, in Aarav's voice.

आज शाम चार बजे, राजेश के साथ आपकी कॉल तय है।

{ "model": "pica-small", "voice": "aarav", "stream": true }

Audio streams back sentence by sentence — each one plays while the next is being made.

Read the speech API

Seven emotions. How it is said matters as much as what is said.

A support agent that sounds sorry, a companion that laughs along, a story that turns frightening — on Pica's expressive lane every sentence can carry the feeling the moment needs.

  • happyJoy that feels real.
  • sadEmotion you can hear.
  • angryHigh energy, on demand.
  • fearTension, naturally.
  • surpriseUnexpected, believably.
  • disgustDistaste, without a word.
  • neutralCalm and clear.

On the expressive lane (Hindi, pica-nano): lane: "expressive" with an emotion.

Everything a voice app needs. At a fraction of the price.

₹21

per hour of audio

Affordable at scale

From ₹21 an hour on pica-nano; ₹32 on pica-small and ₹60 on pica-large, the most natural. Billed by the second.

7

emotions

Expressive

Happy, sad, angry, fear, surprise, disgust or neutral — the same words, said the way the moment needs.

Live

sentence by sentence

Real-time streaming

The first sentence plays while the rest is still being made — no waiting for a whole file.

Yours

designed or cloned

Custom voices & cloning

Describe a voice and get one, or clone one from 10–20 seconds of speech you have the rights to.

12

languages

Multilingual voices

Hindi, English, Marathi, Tamil, Telugu, Gujarati, Bengali, Kannada, Malayalam, Punjabi, Odia and Assamese.

8 kHz

phone-line audio

Built for calls

μ-law frames for a telephone trunk, straight from the stream — for voice agents and IVR.

Benchmarks. As natural as the best, for a fraction of the price.

Pica against the three voice APIs developers compare it with: naturalness on MiniCrow's current benchmark, real-time streaming, and the published list price of an hour of speech.

Swipe the table to compare all four

Benchmark results for Pica, Sarvam Bulbul v3, ElevenLabs and Google Chirp 3 HD: naturalness score, real-time streaming, and list price per hour of speech.
BenchmarkPicaby MiniCrow · nano to large Most affordableSarvamBulbul v3ElevenLabsFlash to v3GoogleChirp 3 HD voices
Naturalness ↑192%88%92%91%
Real-time streaming Yes Yes Yes Yes
List price, per hour of speech2₹21–₹60live price book≈₹151₹30 per 10,000 characters≈₹242–₹484$0.05–$0.10 per 1,000 characters≈₹145$30 per million characters
  1. Naturalness: how human the speech was rated in MiniCrow's current internal benchmark — higher is better.
  2. Published list prices on 24 September 2026. Pica is live from our price book — pica-nano ₹21, pica-small ₹32, pica-large ₹60 an hour of audio. The others charge per character, converted at 14 characters a second (50,400 an hour) and ₹96.1 to the dollar. Check each provider for current rates.

MiniCrow's current benchmark measurements; results vary by language, voice, text, evaluation methodology and audio conditions. Sarvam, ElevenLabs and Google are third-party providers named for comparison only; MiniCrow is not affiliated with them.

Price of an hour of speechRupees, list prices, at 14 characters a second.
Pica
₹21–₹60
Sarvam
₹151
ElevenLabs
₹242–₹484
Google
₹145

Your voice, or a new one. Designed from a sentence, or cloned from a clip.

Design a voiceDescribe it — “warm, unhurried, mid-forties” — and get a new speaker who has never existed. For a brand voice, a narrator or a character.
Clone a voice10–20 seconds of clean speech becomes the voice, for as long as you keep it. You attest that you hold the rights and have the speaker's permission; every attempt is logged and the reference is kept.
Yours aloneA voice belongs to the account that made it: no other customer can list it, speak with it or delete it.

POST /v1/audio/voices

The new voice's id is then just a voice on /v1/audio/speech.

curl

# design a voice from a description
curl https://api.minicrow.com/v1/audio/voices \
  -H "Authorization: Bearer mc_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "Anil", "language": "hi",
       "description": "warm, unhurried, mid-forties"}'

# or clone one from 10–20 seconds of clean speech
curl https://api.minicrow.com/v1/audio/voices \
  -H "Authorization: Bearer mc_YOUR_KEY" \
  -F file=@reference.wav -F name=Anil -F language=hi \
  -F attest=true

Wherever a voice is heard. Pica is speaking.

AI companions
A companion that talks with you in a warm, human voice — and laughs at your jokes.
Audiobooks & learning
Audiobooks, lessons and stories, read the way a person would read them — funny parts included.
  • AI companionsCompanions that don't sound robotic.
  • AI voice agentsSupport, sales, lead qualification, bookings and receptionists.
  • Call centresSpeech, understanding, a response — and a natural voice.
  • Games & AI charactersCharacters with voices of their own, and emotions to match.
  • Content creationYouTube, podcasts, short videos, audiobooks and marketing.
  • E-learningLessons in natural voices, in the learner's language.
  • Enterprise AIInternal support, training and knowledge assistants.
  • IVR & notificationsAlerts, reminders and menus that sound like a person.
  • Brand voicesA voice designed for your brand, or cloned with permission.
  • AccessibilityAnything written, read aloud for whoever needs to hear it.

Illustrative scenes.

Built for voice agents. Hear, understand, act — and answer out loud.

  1. Human
  2. Lark
  3. Osprey
  4. Pica
  5. Human

Lark handles the hearing. Osprey handles the reasoning and the tools. Pica speaks with expression. And Osprey Live brings the real-time experience together, on one phone socket.

A sentence, with feeling

POST /v1/audio/speech — the cost comes back in the headers.

Python

import requests

r = requests.post(
    "https://api.minicrow.com/v1/audio/speech",
    headers={"Authorization": "Bearer mc_YOUR_KEY"},
    json={
        "model": "pica-nano",
        "input": "अरे वाह! आपका ऑर्डर आज ही पहुँच जाएगा।",
        "lane": "expressive",   # Hindi, with emotions
        "emotion": "happy",     # or sad, angry, fear, surprise …
    },
)
open("reply.wav", "wb").write(r.content)
print(r.headers["X-Cost-Paise"], "paise")
  • Stream sentence by sentence

    stream: true sends each sentence as a playable WAV frame the moment it is made.

  • First words sooner

    first_clause speaks the first clause on its own — 437 ms earlier at the median, measured.

  • Straight onto a phone line

    response_format mulaw_8k streams 8 kHz G.711 μ-law frames a trunk takes as they are.

  • Billed by the second

    The playing time of the audio you received — the cost is in the response headers.

Every field, limit and error is in the API reference.

Which Pica is right for you?

Three tiers, one endpoint. Live prices per hour of audio, billed by the second, prepaid in rupees.

The Pica mascot, a glowing magpie, in flight.

 

Pica nano

The expressive one.

Eleven pinned voices, seven emotions on Hindi, and the voices you design or clone.

₹21

per hour of audio · $0.219 an hour

Model id
pica-nano
Voices
11 built-in · your own
Also
Emotions · first_clause · chunk
Best for
Agents · companions · IVR
The Pica mascot landing, its wings raised.

 

Pica small

More voices.

25 Hindi and English voices of its own — and Malayalam, Punjabi, Odia and Assamese.

₹32

per hour of audio · $0.333 an hour

Model id
pica-small
Voices
25 Hindi and English
Also
Streaming · phone audio
Best for
Narration · notifications
The Pica mascot perched on its rock, looking back.

Most natural

Pica large

The most natural.

17 voices, each speaking Hindi and English — and Marathi, Tamil, Telugu, Gujarati, Bengali and Kannada.

₹60

per hour of audio · $0.624 an hour

Model id
pica-large
Voices
17, each in Hindi and English
Also
Streaming · phone audio
Best for
Content · characters · brands
How a request is billed

You pay for the playing time of the audio you received, measured from the audio itself — ₹21 an hour on pica-nano is 0.5833 paise a second. A request that fails before any audio is not charged; a stream you hang up on is charged for the audio made up to that point. The same audio costs the same with or without streaming.

Every MiniCrow price

Voice interfaces should feel less like machines reading text, and more like actual conversations.

  • Speak naturally.
  • Express emotion.
  • React to context.
  • Sound human.

Public launch

1 October 2026.

Build › Deploy › Scale

Real voices. Real emotions. Real conversations. Try every Pica voice free in the Playground today — no account needed.

The Pica mascot, a glowing magpie, singing on its rock.

Questions.

What is Pica?

Pica is MiniCrow's text-to-speech (TTS) API: realistic, expressive and multilingual AI voices, with emotions, real-time streaming and voices of your own — for voice agents, AI companions, call centres, content, e-learning, games and enterprise voice applications.

Which emotions can Pica express?

Seven: happy, sad, angry, fear, surprise, disgust and neutral. Send lane: expressive — the Hindi lane on pica-nano — with an emotion; an emotion the lane doesn't know is refused with a 400, not read flat and billed anyway.

Can I clone a voice or make my own?

Yes. POST /v1/audio/voices designs a new voice from a written description, or clones one from 10–20 seconds of clean speech. Cloning requires attest=true — your statement that you hold the rights and have the speaker's permission — and every attempt is logged. Your voices belong to your account alone and speak on the standard lane.

Does Pica stream in real time?

Yes. With stream: true each sentence arrives as a complete, playable WAV frame as soon as it is made, and first_clause speaks the first clause of the first sentence on its own — 437 ms earlier at the median in our measurement. For phone lines, response_format mulaw_8k streams 8 kHz μ-law.

Which languages does Pica speak?

Hindi and English on every tier. pica-large also speaks Marathi, Tamil, Telugu, Gujarati, Bengali and Kannada; Malayalam, Punjabi, Odia and Assamese are spoken by pica-small. The language follows the script of the text you send.

How much does Pica cost?

₹21 an hour of audio on pica-nano, ₹32 on pica-small and ₹60 on pica-large ($0.219, $0.333 and $0.624) — billed per second of audio received, prepaid in rupees, no subscription. A request that fails before any audio is not charged.

How do I build a voice agent with Pica?

Pica is the voice; you build the harness around it. Hear the caller with Lark, let Osprey decide and call your tools, and speak the reply with Pica, streaming — or use Osprey Live, which runs all three on one phone socket.

When does Pica launch?

MiniCrow launches Pica publicly on 1 October 2026. You can try the voices free in the Playground today.

Meet the MiniCrow family. One API that talks, hears, speaks and watches.

Pica speaks; the rest of MiniCrow thinks, hears and watches. Press play: hear a real call, a recording and its transcript — or watch Lark-V read a real clip.

Osprey

Thinks

Human-like conversational AI — fast, affordable LLMs for assistants, companions and agents, with tool calling.

  • LLM for real-world applications
  • 3 models · 8 modes

Plan my Sunday — something relaxed.

Brunch at 11, a walk by the lake at 4. Shall I book the table?

₹6.9

per million input tokens · $0.072

Lark

Hears

Speech to text for India — Hindi, Marathi, Tamil and more, measured on real 8 kHz phone calls.

  • speech-to-text API for 21 languages
  • 4 tiers · nano · mini · large · Speaker labels
नमस्ते, मेरा ऑर्डर कब आएगा?

Hindi · 8 kHz phone call → text, in real time

₹20

per hour of audio · $0.208

Lark-V

Watches

Video understanding — a summary, a seekable timeline and the spoken words, from a single clip.

  • video summary and timeline API
  • 3 tiers · nano · mini · large
  • 00:04A car pulls into the driveway
  • 00:12Two people unload boxes
  • 00:31Someone waves at the camera

Per clip

billed on what each clip used

Osprey Live

Talks

Real-time multimodal AI that sees, hears, thinks and talks — with tool calling, for AI calling, companions, robots, glasses and cars.

  • real-time multimodal AI model
Live call · 00:42

Caller: “Do you deliver on Sundays?”

Osprey Live: “Yes — until 9 pm. Shall I book you a slot?”

₹65.74

per call-hour, all in · $0.684

Embeddings & Reranking Finds

Semantic search and RAG — embed your documents, rerank a shortlist, and give Osprey your own knowledge to answer from.

Docs