Sees + hears
the picture and the sound
Understands the whole clip
What is on screen and what is said, together — speech quoted in the script it was spoken in.
Introducing
Video understanding for developers: one upload in — a summary, a timeline you can seek and every word spoken out. In Indian languages and code-mix.
Free in the Playground today · Public launch 1 October 2026
Hours of video.Minutes to understand.
Lark-V turns long-form video — YouTube, meetings, lectures, interviews, podcasts, research — into accurate transcripts, summaries, key moments and searchable content. Unstructured video in; useful information out.
Building your own NotebookLM for video, your own HeyPocket or Plaud that also sees, or a searchable library of every meeting? Watching is the part you shouldn't have to build. Lark-V watches the clip, hears the speech and hands your agent JSON — a summary, a timeline and the words spoken — so your time goes into the harness: the product, the prompts, the tools and the memory.
You buildyour code
The harness
# one question about a video
video → Lark-V → summary + timeline
question + that text → Osprey → answer
answer → Pica → voice
We runapi.minicrow.com
The senses
One keyOne billPrepaid in rupeesCost in every response
Pocket, by HeyPocket, and Plaud Note and Plaud NotePin, by Plaud, are pocket-sized AI voice recorders: they capture calls, meetings and conversations and turn them into transcripts, summaries and action items.
Pocket and Plaud hear. Give the device a camera and it becomes more than a recorder — meetings on screen, lectures, site visits, demos. Lark-V watches what was recorded and returns a summary, a timeline and every word spoken; Lark takes the recordings that are sound alone.
Pocket is a product of HeyPocket; Plaud Note and Plaud NotePin are products of Plaud. They are named to describe the kind of device you can build — MiniCrow is not affiliated with them. The device in the videos is an illustration made for this page.
Press play: a real clip, read by lark-v-mini — the timeline lights up as the video reaches each moment.
Hear it in your language
Lark-V
Choose a scenario
Sees the frames, hears the speech · a summary and a timeline you can seek
Lark-V's result
A clip made for this page: two hosts, in Hindi. The summary and the timeline are lark-v-mini's own.
podcast.mp4
Two hosts, one line each, in Hindi
Summary
A man and a woman record a podcast in a studio, discussing voice AI and how AI agents can now converse naturally in Hindi.
A woman in a yellow outfit and a man in a dark blue shirt sit across a table with microphones and headphones in a purple-lit studio with acoustic foam panels. The woman speaks into her mic: "आज के एपिसोड में हम बात करेंगे voice AI की।"
Both hosts smile warmly at each other as the conversation continues.
The man gestures with both hands while speaking: "हां, अब AI agents हिंदी में भी बिल्कुल natural बात करते हैं।"
Sees + hears
the picture and the sound
What is on screen and what is said, together — speech quoted in the script it was spoken in.
0:04
every moment, by the second
Each moment in the clip, stamped with the second it happens — jump your player straight to it.
Every word
with its time
On nano and mini the speech track comes back as a transcript, sentence by sentence with times.
21
languages for the answer
A Tamil video summarised in English, with summary_language — or keep every language as spoken.
3
levels of effort
mid, high or max — from a handful of moments to every distinct one.
Per clip
the cost in every response
Each clip is billed on what it used, and every response shows its cost in paise.
Every moment in order: what is on screen, what is said — in the script it was spoken in — and a summary of the whole.

Like NotebookLM, Google's AI research notebook — it answers from the sources you give it, such as documents, web pages, YouTube videos and audio, and turns them into podcast-style audio overviews.
Refunds are paid within 7 days of the return reaching us.2
Audio overview · 6:12
Build it on MiniCrow
Lark turns audio into text and Lark-V reads the videos. Embeddings & Reranking find the right passage, Osprey answers from it, and Pica reads it aloud.
Apps that watch for you — YouTube and lecture summaries, meeting recaps, a searchable library of every video a team has made.
Build it on MiniCrow
Lark-V turns each video into a summary, a timeline and the words spoken. Embeddings index them, and Osprey answers questions with the moment to jump to.
NotebookLM is a product of Google. They are named only to describe the kind of app you can build — MiniCrow is not affiliated with them. The pictures are illustrations.
Illustrative scenes.
Every tier returns a summary and a timeline you can seek, in any of 21 languages. What changes is how much of the clip it catches.
| Tier | lark-v-nanoThe lightest Lark-V. | lark-v-miniThe one to build on. | lark-v-largeFor video that must be understood. |
|---|---|---|---|
| Accuracy | Good | Better | Best |
| What it catches | The main moments and what is said | More of the detail on screen and in the speech | Every distinct moment, the picture and the sound together |
| Sees what is on screen | Yes | Yes | Yes |
| Hears what is said | Yes, in the script it was spoken in | Yes, in the script it was spoken in | Yes, in the script it was spoken in |
| The words spoken | Word for word, with times | Word for word, with times | Inside the timeline |
| Longest clip | 10 minutes · 20 MB | 10 minutes · 20 MB | 20 MB |
| Best for | Volume · short clips | Summaries · timelines · transcripts | Rich scenes · nuance · long-form |
effort:mid · The key momentshigh · A moment every few seconds — the defaultmax · Every distinct momentOne multipart request with the file. Declare the language spoken, pick the script and the language of the answer, and choose how much detail you want. The answer shown is the real one Lark-V gave for this request, on the clip in the demo above.
The request
POST /v1/video/summaries
Python
import requests
r = requests.post(
"https://api.minicrow.com/v1/video/summaries",
headers={"Authorization": "Bearer mc_YOUR_KEY"},
data={
"model": "lark-v-mini",
"language": "hi", # spoken in the video
"script": "native", # quoted in Devanagari
"summary_language": "en", # answer in English
"effort": "high", # mid · high · max
},
files={"file": open("podcast.mp4", "rb")},
)
video = r.json()
print(video["summary"])
for moment in video["timeline"]:
print(moment["t"], moment["text"])The answer
The first three moments of the timeline, as returned
JSON
{
"model": "lark-v-mini",
"summary": "A man and a woman record a podcast in a studio, discussing voice AI and how AI agents can now converse naturally in Hindi.",
"timeline": [
{
"t": 0,
"text": "A woman in a yellow outfit and a man in a dark blue shirt sit across a table with microphones and headphones in a purple-lit studio with acoustic foam panels. The woman speaks into her mic: \"आज के एपिसोड में हम बात करेंगे voice AI की।\""
},
{
"t": 2,
"text": "Both hosts smile warmly at each other as the conversation continues."
},
{
"t": 4,
"text": "The man gestures with both hands while speaking: \"हां, अब AI agents हिंदी में भी बिल्कुल natural बात करते हैं।\""
}
],
"transcript_segments": [
{
"start": 0,
"end": 2.75,
"text": "आज के एपिसोड में हम बात करेंगे voice AI की।"
},
{
"start": 3.7,
"end": 7.96,
"text": "हां, अब AI agents हिंदी में भी बिल्कुल natural बात करते हैं।"
}
]
}Every field, limit and error is in the API reference.
Three tiers, one endpoint. Each clip is billed on what it used — and every response shows the cost, in paise.

Coming soon
The lightest Lark-V.
The most affordable way to turn a clip into a summary, a timeline and its words — for volume and short clips.
Good accuracy
billed per clip, on what it used

Recommended
The one to build on.
More accurate than nano — for the summaries, timelines and transcripts your product is built on.
Better accuracy
billed per clip, on what it used

For video that must be understood.
Our most accurate — it sees the picture and hears the sound together, and quotes speech in the script it was spoken in.
Best accuracy
billed per clip, on what it used
Each clip is billed on what it used, and every response carries the cost in paise. A silent clip pays nothing for speech.
Public launch
1 October 2026.
Transcribe › Summarize › Understand
Real-world video intelligence, for every developer. Try Lark-V free in the Playground today — no account needed.

Lark-V is MiniCrow's video transcription and understanding API. Send a video and get back a summary, a timeline of what happens and when, and the words spoken — for YouTube videos, meetings, lectures, interviews, podcasts, research and enterprise knowledge, in Indian languages and code-mix.
JSON: a summary, a timeline of moments with the second each happens, and — on lark-v-nano and lark-v-mini — the transcript of the speech track, sentence by sentence with times. Every response carries its cost in paise and says how the clip was watched.
Up to 10 minutes and 20 MB on lark-v-nano and lark-v-mini, and up to 20 MB on lark-v-large. For a longer recording, split it — or send its audio to Lark, which takes up to 25 MB a request.
Declare the language spoken — Hindi, Marathi, Tamil, Telugu, Bengali, Gujarati, Kannada, Malayalam, English, Spanish, German, French, Italian and Arabic are ready — and choose the script: native, romanised or as the model writes it. The summary and timeline can be written in any of 21 languages with summary_language, such as a Tamil video summarised in English.
Accuracy. lark-v-nano, coming soon, will be the lightest and most affordable; lark-v-mini is more accurate; lark-v-large is the most accurate — it sees the picture and hears the sound together. All three return a summary and a timeline you can seek, in any of 21 languages; nano and mini also return the words spoken, sentence by sentence with times.
Each clip is billed on what it used, in paise, and every response shows the cost; a silent clip pays nothing for speech. There is no subscription: you pay for the clips you send, from prepaid rupees.
Yes — that is what it is for. Lark-V turns each video into text your app can search; Embeddings & Reranking find the right moment; Osprey answers with the second to jump to; Pica can read the answer aloud. You build the harness, and MiniCrow handles the watching.
Yes — and give it eyes. Pocket, by HeyPocket, and Plaud Note and Plaud NotePin, by Plaud, are AI voice recorders. Add a camera and Lark-V turns each recording — a meeting on screen, a lecture, a site visit — into a summary, a timeline and the words spoken, while Lark takes the recordings that are sound alone. You build the device and the app; MiniCrow handles the hearing and the watching.
Each timeline moment carries the second it appears. On nano and mini, each sentence of the transcript carries its time too — within about a second on one clear voice.
MiniCrow launches Lark-V publicly on 1 October 2026. You can try it free in the Playground today.
Lark-V watches; the rest of MiniCrow thinks, hears and speaks. Press play: hear a real call, a recording and its transcript, a voice.
Human-like conversational AI — fast, affordable LLMs for assistants, companions and agents, with tool calling.
Plan my Sunday — something relaxed.
Brunch at 11, a walk by the lake at 4. Shall I book the table?
₹6.9
per million input tokens · $0.072
Speech to text for India — Hindi, Marathi, Tamil and more, measured on real 8 kHz phone calls.
Hindi · 8 kHz phone call → text, in real time
₹20
per hour of audio · $0.208
Natural text to speech in Hindi and English, with expressive voices for assistants, IVR and narration.
“Welcome back, Riya! Your order is on its way.”
₹21
per hour of audio · $0.219
Real-time multimodal AI that sees, hears, thinks and talks — with tool calling, for AI calling, companions, robots, glasses and cars.
Caller: “Do you deliver on Sundays?”
Osprey Live: “Yes — until 9 pm. Shall I book you a slot?”
Semantic search and RAG — embed your documents, rerank a shortlist, and give Osprey your own knowledge to answer from.