~150
tokens / sec
Ultra-low latency
Replies land while you're still reading — fast enough for voice and real-time apps.
Introducing
A new generation of fast, affordable LLMs built for AI companions, AI assistants and real-time applications.
Free in the Playground today · Public launch 1 October 2026
Most language models are built to top coding and exam leaderboards. Osprey is built for people. It talks, listens and understands the everyday — your plans, your preferences, your questions, your day — and handles them like a pro.
That makes it the model for what people actually do with AI: talk to it, ask it things, and have it get things done — in companions, assistants, customer support and real-time voice.

Osprey Flash, our flagship — a 552B backbone with 8B prefill and 16B decode — balances intelligence, speed and cost, so an AI that actually feels conversational is practical to build.
~150
tokens / sec
Replies land while you're still reading — fast enough for voice and real-time apps.
48–58
Intelligence Index
The highest score of the four models in our benchmark.
90%
human-like
It sounds like a person — warm, clear and to the point.
Every
mode
Calls your functions and APIs on every Flash and Pro mode.
₹6.9
$0.071 / 1M input
Most affordable in its class. Prepaid in rupees, no subscription.
Every AI can write code.Who takes care of the rest of your life?
The biggest models in the world — Fable, Astra, Kimi — are built for coding and computer use, and they do the rest at 10 to 20 times the cost. Osprey is the LLM built for real life: it connects AI to people, homes, cars, robots and machines.
Plans, preferences, a bad day, a trip, a late order, a table for four. Osprey keeps the thread, answers like a person and uses your tools to finish the job.
Take me to the airport — and avoid tolls.
car.navigate · live map · no tolls
I'll control the car and navigate you in real time — left at the next signal, then the ring road. 42 min, AC at 22°.
Illustrative conversations. The tools are your own functions, called by Osprey.
Osprey is built to operate machines — IoT devices and home appliances, your car and your bike, AR and VR, humanoid robots, farms, drones and air taxis, factories, even defence. It understands people, talks naturally and calls the right tool, fast enough for the real world.
Can you help Dadi with her medicines?
robot.fetch · glass of water
Of course — it's 8 pm, so the blue tablet. I've brought her water too.
Movie night.
home.scene · lights 20% · blinds down
Enjoy! Lights dimmed, blinds down, AC at 23°.
Take me to the airport — avoid tolls.
car.navigate · live
42 minutes, no tolls. I'll guide you and keep the cabin at 22°.
Any fuel on the way?
maps.nearby · fuel
2 km ahead on the left — I'll guide you in your helmet.
What's this building?
Built in 1887, one of the city's oldest — want the story while you walk?
Line 3 is running hot.
plc.set_speed · line 3 · 80%
Spindle at 78 °C — I've slowed it to 80% and alerted maintenance.
Clean the living room before the guests arrive.
vacuum.start · living room · deep clean
On it — a deep clean around the sofa and the rug, done in 25 minutes. I'll dock it before 7.
Which field needs water today?
sensors.soil · field 3 → drone.spray
Field 3 is down to 18% moisture — the drone is watering it now and skipping the rows that are already wet.
Send this parcel to Riya — and get me an air taxi home.
drone.dispatch · parcel → air_taxi.book
The parcel lands on Riya's balcony in 12 minutes, and your air taxi is 4 minutes away.
Illustrative scenes. The tools are your own functions, called by Osprey.
Osprey Flash against three models developers compare it with, on MiniCrow's current benchmark. Best quality result in each row in bold.
Swipe the table to compare all four models
| Benchmark | GPT-5.4 Luna(max) | DeepSeek V4.1 FlashSOTA | Sarvam 105B105B sovereign Indian LLM by Sarvam | |
|---|---|---|---|---|
| Quality | ||||
| Human-like conversation1 | 90% | 78% | 72% | 65% |
| Intelligence Index (AAI)2 | 48–58 | 38 | 40 | 26 |
| Speed and price | ||||
| Output speed, tokens/s | ~150 | ~130 | ~214–266 | ~100 |
| Input price, per 1M tokens3 | $0.115₹11.05 | $0.174₹16.72 | $0.05–$0.444₹4.80–₹42.28 | $0.35₹33.63 |
| Output price, per 1M tokens3 | $0.75₹72.07 | $1.2₹115.32 | $0.14–$1.324₹13.45–₹126.85 | $0.88₹84.57 |
MiniCrow's current benchmark measurements; results vary by workload, prompting, infrastructure and measurement methodology. GPT-5.4 Luna, DeepSeek V4.1 Flash and Sarvam 105B are third-party models named for comparison only; MiniCrow is not affiliated with their makers.
One of the most powerful and affordable LLMs
for making your own JARVIS — an AI companion — even better than GPT-5.4 Luna and DeepSeek V4.1 Flash in our benchmark.
Up to 90% human-like score
Natural conversations. Real understanding. Truly intelligent.
One of the biggest things Osprey is built for is the AI companion — more than just an AI, a companion that truly understands.

Imagine an AI that can:
This opens the door to your own JARVIS-style AI assistant — on an API you can afford to run at scale.
Switch in two lines
Osprey speaks the OpenAI API: change the base URL and the key, and keep the code you have.
Python
from openai import OpenAI
client = OpenAI(
base_url="https://api.minicrow.com/v1",
api_key="mc_YOUR_KEY",
)
reply = client.chat.completions.create(
model="osprey-flash",
messages=[
{"role": "user", "content": "Plan a 5-day Spiti trip for ₹25,000."},
],
)
print(reply.choices[0].message.content)One platform for the whole assistant
Combine Osprey with the rest of MiniCrow — one key, one bill.
A complete real-time AI application stack, through one platform.
Three models, one API. Live prices per million tokens, prepaid in rupees.

Coming soon
The lightest Osprey.
For volume — classification, routing and quick extraction. Reads images, audio and files.
From ₹9.5 per 1M input tokens

Most popular
552B backbone · 8B prefill · 16B decode
The one to build on.
Fast, capable and affordable, with tools on every mode. Right for assistants, companions and agents.
From ₹6.9 per 1M input tokens

For work that must be right.
Long agent loops, long-context reasoning and analysis where a mistake costs more than the tokens.
From ₹79.2 per 1M input tokens
| Model | Mode | For | Input · 1M | Output · 1M |
|---|---|---|---|---|
| Osprey Flash Litecoming soon | speed | Thinking off. For volume. | ₹9.5$0.099 | ₹35.9$0.374 |
| Osprey Flash Litecoming soon | intelligencedefault | Thinking on. The default. | ₹9.5$0.099 | ₹35.9$0.374 |
| Osprey Flash | speed | Short exchanges, low latency. | ₹6.9$0.071 | ₹19.0$0.198 |
| Osprey Flash | intelligencedefault | The default. Tools, images, most work. | ₹7.9$0.082 | ₹26.4$0.275 |
| Osprey Flash | max | When it has to hold up. | ₹21.1$0.220 | ₹126.7$1.32 |
| Osprey Pro | speed | The widest input of any lane. | ₹79.2$0.824 | ₹396.0$4.12 |
| Osprey Pro | intelligencedefault | The default. Deep reasoning. | ₹147.8$1.54 | ₹464.6$4.83 |
| Osprey Pro | max | The top of the catalogue. | ₹211.2$2.20 | ₹1,056.0$10.99 |
Live from our price book; dollars at ₹96.1. Every MiniCrow price
The next generation of AI applications won't be simple chatbots.
They will be AI systems that can see, hear, remember, reason, use tools and communicate naturally. Osprey is our step toward making that infrastructure:
Public launch
1 October 2026.
Build › Deploy › Scale
Build without limits. Talk to Osprey free in the Playground today — no account needed.

Osprey is MiniCrow's family of fast, affordable language models built for human-like conversation — the extrovert of LLMs. It is designed for AI companions and personal AI, assistants, context-aware applications, agents and autonomous workflows, tool and function calling, knowledge-based and real-time applications, customer support and enterprise AI.
Most models are tuned to top coding and exam leaderboards. Osprey is tuned for people: everyday conversation, context, preferences and getting things done with tools. In our benchmark Osprey Flash scored up to 90% on human-like conversation and 48–58 on the Intelligence Index, ahead of GPT-5.4 Luna (78%, 38), DeepSeek V4.1 Flash (72%, 40) and Sarvam 105B (65%, 26), at about 150 tokens per second. Results vary by workload, prompting, infrastructure and measurement methodology.
Osprey is prepaid in rupees with no subscription, and every response carries its own cost. Osprey Flash costs ₹6.9–₹21.1 per million input tokens ($0.071–$0.220) and ₹19.0–₹126.7 per million output tokens ($0.198–$1.32), by mode. The $0.115 and $0.75 in the benchmark are the pricing references used there.
Osprey Flash: 552B backbone / 8B prefill / 16B decode.
Yes. The MiniCrow Playground lets you talk to Osprey free, without an account. API keys open to everyone on 1 October 2026.
Yes. Point the OpenAI SDK at https://api.minicrow.com/v1 with a MiniCrow key and call osprey-flash or osprey-pro (osprey-flash-lite is coming soon). Chat, streaming, tool calling and images use the request shape you already have.
Yes — Osprey Flash and Osprey Pro call tools in every mode. Osprey is one of the best and most effective models for tool calling on everyday, human-level tasks: bookings, orders, calendars, lookups, reminders and whole workflows, done in conversation.
That is one of the biggest use cases Osprey is built for: a companion that remembers your context, understands your preferences, has natural conversations, uses tools on your behalf, plans tasks, calls APIs and executes workflows — quickly enough to feel conversational. Add Lark to hear, Pica to speak and Osprey Live for real-time calls.
MiniCrow launches Osprey publicly on 1 October 2026. You can talk to it in the Playground today.
Osprey thinks; the rest of MiniCrow gives it a voice, ears and eyes. Press play: hear a real call, a recording and its transcript, a voice — or watch Lark-V read a real clip.
Speech to text for India — Hindi, Marathi, Tamil and more, measured on real 8 kHz phone calls.
Hindi · 8 kHz phone call → text, in real time
₹20
per hour of audio · $0.208
Natural text to speech in Hindi and English, with expressive voices for assistants, IVR and narration.
“Welcome back, Riya! Your order is on its way.”
₹21
per hour of audio · $0.219
Video understanding — a summary, a seekable timeline and the spoken words, from a single clip.
Per clip
billed on what each clip used
Real-time multimodal AI that sees, hears, thinks and talks — with tool calling, for AI calling, companions, robots, glasses and cars.
Caller: “Do you deliver on Sundays?”
Osprey Live: “Yes — until 9 pm. Shall I book you a slot?”
Semantic search and RAG — embed your documents, rerank a shortlist, and give Osprey your own knowledge to answer from.