Introducing

Osprey LiveIndia's only real LLM model with vision and audio input.

A real-time, multimodal AI model that understands the world like humans — see, hear, speak and reason. Built for the next generation of AI products.

  • Real-time & low latency
  • Optimized for Indian languages
  • Custom tool calling
  • Advanced reasoning
  • Built for developers

Multimodal · Real-time · Open for builders · Public launch 1 October 2026

The next generation of AI applications won't just read text.They will see, hear, understand, reason, speak and act — in real time.

  1. See
  2. Hear
  3. Understand
  4. Reason
  5. Act
  6. Speak

Meet Osprey Live by MiniCrow AI — a real-time, multimodal LLM built for developers creating AI systems that interact with the physical and digital world. Vision, audio, text, real-time reasoning and tool calling, together in a single AI experience.

What is Osprey Live? One real-time model, across six modalities.

Osprey Live is designed for applications where AI needs to understand more than a text prompt. Developers can build AI systems that listen, see, reason and respond, instead of relying on separate models for every interaction — think of it as a real-time AI brain for applications that need to interact naturally with humans and the environment.

  • Audio → AudioReal-time voice conversations
  • Text → AudioNatural & expressive speech
  • Text → TextSmart, contextual responses
  • Audio → TextAccurate transcription
  • Video → TextUnderstand video content
  • Video → AudioTalk about what you see

One model. Multiple modalities.

Instead of stitching together separate models for speech to text, the LLM, vision, text to speech and tool calling, Osprey Live brings the real-time intelligence layer together for developers.

The usual way

Developers traditionally have to combine

  • Vision model+
  • STT+
  • LLM+
  • TTS+
  • Agent framework+
  • Tool layer+
  • Memory

Seven services to wire, pay for, keep in sync — and wait on, one after another.

Osprey Live

  1. Audio + Vision + Text
  2. Reasoning
  3. Tool Calling + Memory
  4. Audio / Text / Multimodal Response

Less infrastructure to stitch together. More time to build the actual product.

Made for real life. It sees, hears and gets things done — live.

Glasses that tell you what you're looking at, a car that finds the parking, a front desk that reads the badge, a clinic line that books the patient in. Osprey Live keeps the thread, answers out loud and uses your tools to finish the job.

9:41
Osprey LiveHome assistant · video
LIVESees the strip

Is this the one I take after lunch?

medication.schedule · after lunch

Yes — that's your after-lunch tablet. One, with water. I'll remind you about the evening one at 9.

Illustrative conversations. The tools are your own functions, called by Osprey Live.

Sees. Hears. Thinks. And uses your tools.

Vision

Give your AI the ability to understand:

  • Images
  • Video
  • Visual environments
  • Objects and scenes
  • Real-world context

This opens the door to AI systems that can understand what the user is actually looking at.

The Osprey Live mascot's head, in dark sunglasses.

Audio understanding

Osprey Live processes audio input for conversational interactions — so users can speak naturally instead of typing everything.

Reasoning

Designed for complex conversational reasoning, contextual understanding and multi-step interactions.

57intelligence & reasoning score

Custom tool calling

Connect your own tools, APIs, databases and application logic. Your AI can move beyond answering questions and interact with the systems around it.

  1. User
  2. AI
  3. Tool
  4. API
  5. Result
  6. AI
  7. Response

This enables agentic workflows and real-world automation.

Hear it. A real-time conversation, tools and all.

Press play: an illustrative call. Osprey Live hears the caller, checks your systems with a tool call and answers out loud, in real time.

Hear it in your language

Osprey Live

Choose a scenario

₹67–₹110 / hour, all inYour tools, mid-callReal time
Call in progress

Live call · hears, thinks and speaks on one WebSocket

Illustrative call

Spoken by Pica: Tanvi is the agent, Aarav the caller. The conversation is illustrative.

नमस्ते! क्लिनिक से बोल रही हूँ। आपको किस दिन का अपॉइंटमेंट चाहिए?

कल शाम पाँच बजे के बाद मिल जाएगा?

check_appointment_slots · tomorrow, after 17:00

कल साढ़े पाँच बजे का स्लॉट खाली है। बुक कर दूँ?

हाँ, कर दीजिए।

send_sms_confirmation

हो गया — कन्फ़र्मेशन SMS भेज दिया है।

Read the voice agent API

Real-time AI. A different architecture.

Real-time interaction changes the architecture of AI applications.

Instead of

  1. User
  2. Request
  3. Wait
  4. Response

Osprey Live is designed around

  1. Listen
  2. Understand
  3. Reason
  4. Act
  5. Respond

with streaming and low-latency interaction.

The Osprey Live mascot gliding at speed, trailing light.

2.8–4.1 s

to the first response, including tool calls and memory recall

First response latencySeconds, including tool calls and memory recall — lower is better.
Osprey Live
2.8–4.1 s
Gemini 3.1 Flash Live
3.5–6 s
Sarvam Stack
4–7 s

These are MiniCrow's current benchmark measurements and can vary based on workload, network conditions, tools, memory operations and evaluation setup.

Osprey Live vs real-time AI models. For voice, vision and beyond.

Osprey Live against Google Gemini 3.1 Flash Live and the Sarvam Stack (Sarvam + Saaras + Bulbul), across pricing, multimodal capabilities, reasoning, voice and latency.

Swipe the table to compare all three

Osprey Live compared with Google Gemini 3.1 Flash Live and the Sarvam Stack: price per hour, video input, intelligence and reasoning score, hearing and TTS score, voice emotions, first response latency and supported modalities.
FeatureOsprey LiveReal-time. Multimodal. Built for what's next. Most capableGoogle Gemini 3.1 Flash LiveGoogle's most advanced multimodal model.Sarvam StackSarvam + Saaras + Bulbul · India's sovereign AI stack.
Price, per hour1₹67–₹110≈₹170–₹210≈₹135–₹190
Video inputYesImage + video, real timeYesImage + video, real timeNoAudio only
Intelligence & reasoning score (out of 100)2575349
Hearing / TTS score (out of 100)2949095
AI voice with emotionsYesUp to 7 emotionsNoLimited expressivenessNoNeutral voices
First response latency3incl. tool calls & memory recall2.8–4.1 s3.5–6 s4–7 s
Supported modalities
  • Voice
  • Image
  • Video
  • Text
  • Voice
  • Image
  • Video
  • Text
  • Voice
  • Text
In shortMost capable. Most versatile. Best for real-world AI.Powerful and general purpose. Great for multimodal tasks.Optimized for Indian languages. Strong in voice, limited modality.
  1. Approximate price of an hour of real-time use, from MiniCrow's current comparison.
  2. The scores shown above are MiniCrow's benchmark measurements, not universal industry rankings. Results can vary by benchmark design, prompts, languages, model versions and infrastructure.
  3. MiniCrow's current benchmark measurements; they can vary based on workload, network conditions, tools, memory operations and evaluation setup.

Google and Sarvam are third-party providers named for comparison only; MiniCrow is not affiliated with them.

Real-time AI without extreme infrastructure costs.

Real-time multimodal AI can become expensive very quickly. For developers building applications that run thousands of hours of real-time AI interactions, inference economics can have a major impact on product viability — Osprey Live is designed around making real-time multimodal intelligence accessible to developers and startups.

Osprey Live

₹67–₹110

per hour · $0.70–$1.14 · prepaid in rupees, no subscription

Model id
osprey-live
Endpoint
wss://api.minicrow.com/v1/agent/live
Billed
On what each session uses — the cost comes back on every turn
Price of an hour of real-time AIRupees, approximate — MiniCrow's current comparison.
Osprey Live
₹67–₹110
Gemini 3.1 Flash Live
₹170–₹210
Sarvam Stack
₹135–₹190
  • Osprey Live

    ₹67–₹110

    per hour · $0.70–$1.14

  • Gemini 3.1 Flash Live

    ≈₹170–₹210

    per hour · $1.77–$2.19

  • Sarvam Stack

    ≈₹135–₹190

    per hour · $1.40–$1.98

AI voice with emotion. Same emotions. Different possibilities.

Traditional AI voice systems often focus primarily on generating understandable speech. Osprey Live goes further by supporting expressive AI voice interactions — up to seven emotions.

  • Neutral
  • Happy
  • Sad
  • Angry
  • Fear
  • Surprise
  • Disgust

Especially useful for

  • AI companions
  • Voice assistants
  • Customer support
  • Interactive characters
  • AI calling
  • Education
  • Entertainment

The goal is not simply:

AI speaks.

It's:

AI understands → reasons → responds naturally.

Built for the physical world. AI outside the browser.

Osprey Live becomes particularly interesting when AI moves outside the browser — into glasses, robots, cars and phone lines.

AI-enabled VR glasses
What am I looking at?This is India Gate, a war memorial in New Delhi.

Build assistants that can:

  • See what the user sees
  • Listen to the user
  • Understand the environment
  • Answer questions
  • Provide contextual assistance
AI humanoid robots

Give robots a multimodal conversational intelligence layer capable of:

  • Seeing
  • Hearing
  • Reasoning
  • Speaking
  • Tool calling
AI-enabled vehicles
Navigate to Connaught Place

Build intelligent in-car assistants that can:

  • Understand voice commands
  • See visual information
  • Navigate
  • Interact with vehicle systems
  • Provide contextual assistance
Smart AI tele-calling bots

Create AI agents capable of handling natural conversations for:

  • Customer support
  • Sales
  • Appointment booking
  • Lead qualification
  • Business automation
AI healthcare assistants

Build conversational assistants capable of understanding voice and visual context while connecting to approved healthcare workflows and tools.

AI education

Create interactive tutors that can:

  • Listen to students
  • Understand visual material
  • Explain concepts
  • Conduct conversations
  • Provide personalized assistance

Illustrative scenes.

Build your own AI companion.

One of the most interesting applications for Osprey Live is the creation of personal AI companions. Imagine an AI that can:

  • See what you see
  • Hear what you say
  • Remember context
  • Use your tools
  • Have natural conversations
  • Respond with expressive voice

Instead of a chatbot that lives inside a text box, developers can build AI that feels much more integrated into everyday life.

The Osprey Live mascot in flight, wings wide, in dark sunglasses.

Built for what's next. Real intelligence. Real impact.

  • AI assistantsAssistants that see, hear and act for the people who use them.
  • AI companionsCompanions that remember, and answer in an expressive voice.
  • Customer supportSupport that listens, checks your systems and answers live.
  • AI tele-callingSales, bookings and lead qualification over the phone.
  • Meeting assistantsAn assistant in the meeting that follows along and answers.
  • Real-time translationSpeak in one language and be understood in another.
  • Education & learningTutors that listen, look at the material and explain.
  • Healthcare assistantsVoice and visual context, connected to approved workflows.
  • Content creationMultimodal content applications that talk about what they see.
  • And lots moreVR glasses, humanoid robots, vehicles and intelligent devices.

Built for developers. Make real-time intelligence an API.

You build the harness — the product, the prompts, the tools and the memory. Osprey Live is the real-time intelligence layer under it.

You buildyour code

The harness

  • Your productThe app your users open — on the web, on a phone, on a device.
  • Prompts and personalityHow it thinks, what it says, how it sounds.
  • Tools and workflowsYour APIs and the actions it is allowed to take.
  • Memory and knowledgeWhat it remembers, and the sources it answers from.
  • EvalsHow you know it works, call after call.

# one turn of a real-time app

voice + camera → Osprey Live → reply + tool call

your tool's result → Osprey Live → spoken answer

A real-time session, in Python

GET wss://api.minicrow.com/v1/agent/live — speaker, stream_microphone and run_tool are your own code.

Python

# pip install "websockets>=14"
import asyncio, json, os
import websockets

URL = ("wss://api.minicrow.com/v1/agent/live?model=osprey-live&language=hi"
       "&encoding=linear16&sample_rate=16000&output=pcm16_24k&voice=david")
KEY = {"Authorization": f"Bearer {os.environ['MINICROW_API_KEY']}"}
AGENT = json.load(open("agent.json"))   # who it is, your brief, your tools

async def main():
    async with websockets.connect(URL, additional_headers=KEY) as ws:
        async for message in ws:
            if isinstance(message, bytes):          # the reply, spoken
                speaker.play(message)
                continue
            event = json.loads(message)
            if event["type"] == "session.started":
                await ws.send(json.dumps({"type": "session.configure", "agent": AGENT}))
            elif event["type"] == "session.configured":
                asyncio.create_task(stream_microphone(ws))
            elif event["type"] == "tool.call":      # your code, your systems
                output = await run_tool(event["name"], event["arguments"])
                await ws.send(json.dumps({"type": "tool.result",
                                          "call_id": event["call_id"], "output": output}))
            elif event["type"] == "turn.usage":
                print(f"turn {event['turn']}: {event['cost']} paise")

asyncio.run(main())
  • One WebSocket

    Stream the user's voice in; the reply comes back spoken, while it is still being written.

  • Your tools, mid-conversation

    tool.call arrives with checked arguments; answer with tool.result and it speaks the outcome.

  • Interruptions handled

    Talk over it and it stops — or carries on after a cough or an “hmm”.

  • Cost on every turn

    turn.usage says what each turn cost, in paise, from a prepaid balance in rupees.

Osprey Live is designed for developers building

  • Real-time AI applications
  • AI agents
  • Voice AI
  • Multimodal AI
  • AI companions
  • AI assistants
  • Robotics
  • VR applications
  • Smart vehicles
  • AI calling
  • Customer support automation
  • Interactive AI
  • Enterprise AI
  • Intelligent devices

Every field, limit and error is in the API reference.

Optimized for Indian applications. Designed for India, built for the world.

Osprey Live is designed with Indian-language and multilingual applications in mind. While being designed for Indian use cases, the underlying architecture is intended for global real-time AI applications.

  • Indian voice assistants
  • Multilingual AI agents
  • Regional-language applications
  • AI calling systems
  • Customer support automation
  • Voice-first applications

The future of AI is multimodal

From AI companions and VR glasses to humanoid robots, intelligent vehicles and AI calling systems, developers can now build applications where AI interacts with the world in real time.

See. Hear. Think. Talk. Build.

Launching 1 October 2026

The Osprey Live mascot diving head-first, wings spread.

Questions.

What is Osprey Live?

Osprey Live by MiniCrow is a real-time, multimodal LLM built for developers creating AI systems that interact with the physical and digital world. It brings vision, audio, text, real-time reasoning and tool calling together into a single AI experience: it can see, hear, understand, reason, speak and act — in real time.

What can Osprey Live see and hear?

It processes audio input for natural, spoken conversation, and understands images, video, visual environments, objects and scenes — so an AI can understand what the user is actually looking at, and talk about it.

Which input and output modalities does Osprey Live support?

Audio → audio (real-time voice conversations), text → audio (natural, expressive speech), text → text (smart, contextual responses), audio → text (accurate transcription), video → text (understand video content) and video → audio (talk about what you see).

Can Osprey Live call my own tools and APIs?

Yes. Connect your own tools, APIs, databases and application logic. The model calls them mid-conversation, your code returns the result, and it answers with it: User → AI → Tool → API → Result → AI → Response. That is what makes agentic workflows and real-world automation possible.

How fast is Osprey Live?

In MiniCrow's current benchmark, 2.8–4.1 s to the first response, including tool calls and memory recall — against 3.5–6 s for Gemini 3.1 Flash Live and 4–7 s for the Sarvam Stack. It can vary with workload, network conditions, tools, memory operations and evaluation setup.

How much does Osprey Live cost?

About ₹67–₹110 an hour ($0.70–$1.14), against ₹170–₹210 for Gemini 3.1 Flash Live and ₹135–₹190 for the Sarvam Stack in MiniCrow's current comparison. Prepaid in rupees, with no subscription.

Can Osprey Live's voice express emotion?

Yes — up to seven emotions: neutral, happy, sad, angry, fear, surprise and disgust. It is built for expressive voice interactions: AI companions, voice assistants, customer support, interactive characters, AI calling, education and entertainment.

Is Osprey Live built for Indian languages?

Yes. It is designed with Indian-language and multilingual applications in mind — Indian voice assistants, multilingual AI agents, regional-language applications, AI calling and voice-first apps — while the underlying architecture is intended for global real-time AI applications.

How is Osprey Live different from Gemini 3.1 Flash Live or the Sarvam Stack?

In MiniCrow's current comparison, Osprey Live scores 57 on intelligence and reasoning (Gemini 3.1 Flash Live 53, the Sarvam Stack 49), takes image and video input in real time, speaks with up to 7 emotions, answers first in 2.8–4.1 s and costs ₹67–₹110 an hour. The Sarvam Stack is audio only; Gemini 3.1 Flash Live costs ₹170–₹210 an hour.

When does Osprey Live launch?

MiniCrow launches Osprey Live publicly on 1 October 2026. The API reference is open to read today.

Meet the MiniCrow family. One API that talks, hears, speaks and watches.

Osprey Live is part of the broader MiniCrow AI platform: Osprey for LLM reasoning and conversation, Lark for speech to text, Pica for expressive AI voice, Lark-V for video understanding, and embeddings and reranking for search and RAG — the building blocks for real-time AI agents and intelligent applications.

Osprey

Thinks

Human-like conversational AI — fast, affordable LLMs for assistants, companions and agents, with tool calling.

  • LLM for real-world applications
  • 3 models · 8 modes

Plan my Sunday — something relaxed.

Brunch at 11, a walk by the lake at 4. Shall I book the table?

₹6.9

per million input tokens · $0.072

Lark

Hears

Speech to text for India — Hindi, Marathi, Tamil and more, measured on real 8 kHz phone calls.

  • speech-to-text API for 21 languages
  • 4 tiers · nano · mini · large · Speaker labels
नमस्ते, मेरा ऑर्डर कब आएगा?

Hindi · 8 kHz phone call → text, in real time

₹20

per hour of audio · $0.208

Pica

Speaks

Natural text to speech in Hindi and English, with expressive voices for assistants, IVR and narration.

  • text-to-speech API for Hindi and English
  • 3 tiers · nano · small · large

“Welcome back, Riya! Your order is on its way.”

₹21

per hour of audio · $0.219

Lark-V

Watches

Video understanding — a summary, a seekable timeline and the spoken words, from a single clip.

  • video summary and timeline API
  • 3 tiers · nano · mini · large
  • 00:04A car pulls into the driveway
  • 00:12Two people unload boxes
  • 00:31Someone waves at the camera

Per clip

billed on what each clip used

Embeddings & Reranking Finds

Semantic search and RAG — embed your documents, rerank a shortlist, and give Osprey your own knowledge to answer from.

Docs