Introducing

Lark-VTranscribe. Summarize. Understand.

Video understanding for developers: one upload in — a summary, a timeline you can seek and every word spoken out. In Indian languages and code-mix.

Free in the Playground today · Public launch 1 October 2026

Hours of video.Minutes to understand.

Lark-V turns long-form video — YouTube, meetings, lectures, interviews, podcasts, research — into accurate transcripts, summaries, key moments and searchable content. Unstructured video in; useful information out.

You build the harness. We handle the watching.

Building your own NotebookLM for video, your own HeyPocket or Plaud that also sees, or a searchable library of every meeting? Watching is the part you shouldn't have to build. Lark-V watches the clip, hears the speech and hands your agent JSON — a summary, a timeline and the words spoken — so your time goes into the harness: the product, the prompts, the tools and the memory.

You buildyour code

The harness

  • Your productThe app your users open — on the web, on a phone, on a device.
  • Prompts and personalityHow it thinks, what it says, how it sounds.
  • Tools and workflowsYour APIs and the actions it is allowed to take.
  • Memory and knowledgeWhat it remembers, and the sources it answers from.
  • EvalsHow you know it works, call after call.

# one question about a video

video → Lark-V → summary + timeline

question + that text → Osprey → answer

answer → Pica → voice

Build your own HeyPocket or Plaud. Now with eyes.

Pocket, by HeyPocket, and Plaud Note and Plaud NotePin, by Plaud, are pocket-sized AI voice recorders: they capture calls, meetings and conversations and turn them into transcripts, summaries and action items.

Pocket and Plaud hear. Give the device a camera and it becomes more than a recorder — meetings on screen, lectures, site visits, demos. Lark-V watches what was recorded and returns a summary, a timeline and every word spoken; Lark takes the recordings that are sound alone.

On the table
A meeting, every speaker in the transcript.
On the phone
Clipped to the back of a phone: the whole call, both sides.
On you
A pin with a camera: what you saw, and what was said.
  1. A recording with a picture
  2. Lark-V — a summary, a timeline and the words
  3. Osprey — answers, with the moment to jump to
  4. Embeddings — every moment, searchable
  5. Your app

Pocket is a product of HeyPocket; Plaud Note and Plaud NotePin are products of Plaud. They are named to describe the kind of device you can build — MiniCrow is not affiliated with them. The device in the videos is an illustration made for this page.

One upload. A summary, a timeline and every word.

Press play: a real clip, read by lark-v-mini — the timeline lights up as the video reaches each moment.

Open the Playground

Hear it in your language

Lark-V

Choose a scenario

Summary + timelineSpeech in its own scriptCost in every response

Sees the frames, hears the speech · a summary and a timeline you can seek

Lark-V's result

A clip made for this page: two hosts, in Hindi. The summary and the timeline are lark-v-mini's own.

podcast.mp4

Two hosts, one line each, in Hindi

0:08

Summary

A man and a woman record a podcast in a studio, discussing voice AI and how AI agents can now converse naturally in Hindi.

0:00

A woman in a yellow outfit and a man in a dark blue shirt sit across a table with microphones and headphones in a purple-lit studio with acoustic foam panels. The woman speaks into her mic: "आज के एपिसोड में हम बात करेंगे voice AI की।"

0:02

Both hosts smile warmly at each other as the conversation continues.

0:04

The man gestures with both hands while speaking: "हां, अब AI agents हिंदी में भी बिल्कुल natural बात करते हैं।"

lark-v-mini · language=hi · script=native · summary_language=hi
Read the video API

Built to understand video. Not just to caption it.

Sees + hears

the picture and the sound

Understands the whole clip

What is on screen and what is said, together — speech quoted in the script it was spoken in.

0:04

every moment, by the second

A seekable timeline

Each moment in the clip, stamped with the second it happens — jump your player straight to it.

Every word

with its time

The transcript, too

On nano and mini the speech track comes back as a transcript, sentence by sentence with times.

21

languages for the answer

Summaries in any language

A Tamil video summarised in English, with summary_language — or keep every language as spoken.

3

levels of effort

As detailed as you need

mid, high or max — from a handful of moments to every distinct one.

Per clip

the cost in every response

Pay for what it watched

Each clip is billed on what it used, and every response shows its cost in paise.

Sees what happens. Hears what's said.

Every moment in order: what is on screen, what is said — in the script it was spoken in — and a summary of the whole.

The Lark-V mascot swooping down, watching.
  • 0:00 · seenTwo hosts at a table in a purple-lit studio, microphones on.
  • 0:00 · heardआज के एपिसोड में हम बात करेंगे voice AI की।
  • 0:04 · seenHe gestures with both hands as he speaks.
  • 0:04 · heardहां, अब AI agents हिंदी में भी बिल्कुल natural बात करते हैं।
  • SummaryA podcast on how AI agents now talk naturally in Hindi.

Build your own. A NotebookLM for video, a video assistant of your own.

Your own NotebookLM

Like NotebookLM, Google's AI research notebook — it answers from the sources you give it, such as documents, web pages, YouTube videos and audio, and turns them into podcast-style audio overviews.

  • Policy.pdf
  • Webinar
  • Call notes
  • Help page

Refunds are paid within 7 days of the return reaching us.2

Audio overview · 6:12

Build it on MiniCrow

Lark turns audio into text and Lark-V reads the videos. Embeddings & Reranking find the right passage, Osprey answers from it, and Pica reads it aloud.

  • Lark
  • Lark-V
  • Embeddings & Reranking
  • Osprey
  • Pica

Your own video assistant

Apps that watch for you — YouTube and lecture summaries, meeting recaps, a searchable library of every video a team has made.

  • 00:00Introduction
  • 03:12Pricing, explained
  • 07:40Questions from the audience

Build it on MiniCrow

Lark-V turns each video into a summary, a timeline and the words spoken. Embeddings index them, and Osprey answers questions with the moment to jump to.

  • Lark-V
  • Embeddings & Reranking
  • Osprey

NotebookLM is a product of Google. They are named only to describe the kind of app you can build — MiniCrow is not affiliated with them. The pictures are illustrations.

Every video, understood. From lecture halls to YouTube.

Lectures & e-learning
A lecture becomes notes, chapters and every word said — in the language it was taught.
Creators & YouTube
A long video, summarised with a timeline of its moments — ready for chapters and search.
  • YouTube summariesSummaries and chapters for long videos.
  • Meetings & note takingRecaps, decisions and what was said, when.
  • E-learningLectures turned into notes, chapters and study material.
  • Interviews & podcastsEvery word, and the moments worth clipping.
  • ResearchHours of footage, searchable in minutes.
  • Enterprise knowledgeA searchable library of every video your team has made.
  • Screen recordings & demosWalkthroughs and product demos, step by step.
  • Ad & creative reviewWhat is on screen and what is said, second by second.
  • Games & charactersScenes and gameplay described, for characters that react.
  • AccessibilityDescriptions and transcripts for anyone who can't watch or hear.

Illustrative scenes.

Three tiers. Pick by how much you need to catch.

Every tier returns a summary and a timeline you can seek, in any of 21 languages. What changes is how much of the clip it catches.

The three Lark-V tiers compared: accuracy, what each catches, the transcript, the longest clip and what each is best for.
Tierlark-v-nanoThe lightest Lark-V.lark-v-miniThe one to build on.lark-v-largeFor video that must be understood.
Accuracy Good Better Best
What it catchesThe main moments and what is saidMore of the detail on screen and in the speechEvery distinct moment, the picture and the sound together
Sees what is on screenYesYesYes
Hears what is saidYes, in the script it was spoken inYes, in the script it was spoken inYes, in the script it was spoken in
The words spokenWord for word, with timesWord for word, with timesInside the timeline
Longest clip10 minutes · 20 MB10 minutes · 20 MB20 MB
Best forVolume · short clipsSummaries · timelines · transcriptsRich scenes · nuance · long-form
Detail, on every tier, with effort:mid · The key momentshigh · A moment every few seconds — the defaultmax · Every distinct moment

Built for developers. Structured JSON, not a wall of text.

One multipart request with the file. Declare the language spoken, pick the script and the language of the answer, and choose how much detail you want. The answer shown is the real one Lark-V gave for this request, on the clip in the demo above.

The request

POST /v1/video/summaries

Python

import requests

r = requests.post(
    "https://api.minicrow.com/v1/video/summaries",
    headers={"Authorization": "Bearer mc_YOUR_KEY"},
    data={
        "model": "lark-v-mini",
        "language": "hi",          # spoken in the video
        "script": "native",        # quoted in Devanagari
        "summary_language": "en",  # answer in English
        "effort": "high",          # mid · high · max
    },
    files={"file": open("podcast.mp4", "rb")},
)
video = r.json()
print(video["summary"])
for moment in video["timeline"]:
    print(moment["t"], moment["text"])

The answer

The first three moments of the timeline, as returned

JSON

{
  "model": "lark-v-mini",
  "summary": "A man and a woman record a podcast in a studio, discussing voice AI and how AI agents can now converse naturally in Hindi.",
  "timeline": [
    {
      "t": 0,
      "text": "A woman in a yellow outfit and a man in a dark blue shirt sit across a table with microphones and headphones in a purple-lit studio with acoustic foam panels. The woman speaks into her mic: \"आज के एपिसोड में हम बात करेंगे voice AI की।\""
    },
    {
      "t": 2,
      "text": "Both hosts smile warmly at each other as the conversation continues."
    },
    {
      "t": 4,
      "text": "The man gestures with both hands while speaking: \"हां, अब AI agents हिंदी में भी बिल्कुल natural बात करते हैं।\""
    }
  ],
  "transcript_segments": [
    {
      "start": 0,
      "end": 2.75,
      "text": "आज के एपिसोड में हम बात करेंगे voice AI की।"
    },
    {
      "start": 3.7,
      "end": 7.96,
      "text": "हां, अब AI agents हिंदी में भी बिल्कुल natural बात करते हैं।"
    }
  ]
}

Every field, limit and error is in the API reference.

Which Lark-V is right for you?

Three tiers, one endpoint. Each clip is billed on what it used — and every response shows the cost, in paise.

The Lark-V mascot, a glowing violet lark, swooping down to watch.

Coming soon

Lark-V nano

The lightest Lark-V.

The most affordable way to turn a clip into a summary, a timeline and its words — for volume and short clips.

Good accuracy

billed per clip, on what it used

Coming soonDocs
Model id
lark-v-nano
The words spoken
Word for word, with times
Limit
10 minutes · 20 MB
Best for
Volume · short clips
The Lark-V mascot, a glowing violet lark, face on with its wings raised.

Recommended

Lark-V mini

The one to build on.

More accurate than nano — for the summaries, timelines and transcripts your product is built on.

Better accuracy

billed per clip, on what it used

Model id
lark-v-mini
The words spoken
Word for word, with times
Limit
10 minutes · 20 MB
Best for
Summaries · timelines · transcripts
The Lark-V mascot seen from above, wings spread.

 

Lark-V large

For video that must be understood.

Our most accurate — it sees the picture and hears the sound together, and quotes speech in the script it was spoken in.

Best accuracy

billed per clip, on what it used

Model id
lark-v-large
The words spoken
Inside the timeline
Limit
20 MB
Best for
Rich scenes · nuance · long-form
How a clip is billed

Each clip is billed on what it used, and every response carries the cost in paise. A silent clip pays nothing for speech.

Every MiniCrow price

Public launch

1 October 2026.

Transcribe › Summarize › Understand

Real-world video intelligence, for every developer. Try Lark-V free in the Playground today — no account needed.

Two Lark-V mascots in flight, one singing to the other.

Questions.

What is Lark-V?

Lark-V is MiniCrow's video transcription and understanding API. Send a video and get back a summary, a timeline of what happens and when, and the words spoken — for YouTube videos, meetings, lectures, interviews, podcasts, research and enterprise knowledge, in Indian languages and code-mix.

What does Lark-V return?

JSON: a summary, a timeline of moments with the second each happens, and — on lark-v-nano and lark-v-mini — the transcript of the speech track, sentence by sentence with times. Every response carries its cost in paise and says how the clip was watched.

How long can a video be?

Up to 10 minutes and 20 MB on lark-v-nano and lark-v-mini, and up to 20 MB on lark-v-large. For a longer recording, split it — or send its audio to Lark, which takes up to 25 MB a request.

Which languages does Lark-V understand?

Declare the language spoken — Hindi, Marathi, Tamil, Telugu, Bengali, Gujarati, Kannada, Malayalam, English, Spanish, German, French, Italian and Arabic are ready — and choose the script: native, romanised or as the model writes it. The summary and timeline can be written in any of 21 languages with summary_language, such as a Tamil video summarised in English.

What is the difference between the tiers?

Accuracy. lark-v-nano, coming soon, will be the lightest and most affordable; lark-v-mini is more accurate; lark-v-large is the most accurate — it sees the picture and hears the sound together. All three return a summary and a timeline you can seek, in any of 21 languages; nano and mini also return the words spoken, sentence by sentence with times.

How much does Lark-V cost?

Each clip is billed on what it used, in paise, and every response shows the cost; a silent clip pays nothing for speech. There is no subscription: you pay for the clips you send, from prepaid rupees.

Can I build my own NotebookLM for video with Lark-V?

Yes — that is what it is for. Lark-V turns each video into text your app can search; Embeddings & Reranking find the right moment; Osprey answers with the second to jump to; Pica can read the answer aloud. You build the harness, and MiniCrow handles the watching.

Can I build my own HeyPocket or Plaud with Lark-V?

Yes — and give it eyes. Pocket, by HeyPocket, and Plaud Note and Plaud NotePin, by Plaud, are AI voice recorders. Add a camera and Lark-V turns each recording — a meeting on screen, a lecture, a site visit — into a summary, a timeline and the words spoken, while Lark takes the recordings that are sound alone. You build the device and the app; MiniCrow handles the hearing and the watching.

How exact are the times?

Each timeline moment carries the second it appears. On nano and mini, each sentence of the transcript carries its time too — within about a second on one clear voice.

When does Lark-V launch?

MiniCrow launches Lark-V publicly on 1 October 2026. You can try it free in the Playground today.

Meet the MiniCrow family. One API that talks, hears, speaks and watches.

Lark-V watches; the rest of MiniCrow thinks, hears and speaks. Press play: hear a real call, a recording and its transcript, a voice.

Osprey

Thinks

Human-like conversational AI — fast, affordable LLMs for assistants, companions and agents, with tool calling.

  • LLM for real-world applications
  • 3 models · 8 modes

Plan my Sunday — something relaxed.

Brunch at 11, a walk by the lake at 4. Shall I book the table?

₹6.9

per million input tokens · $0.072

Lark

Hears

Speech to text for India — Hindi, Marathi, Tamil and more, measured on real 8 kHz phone calls.

  • speech-to-text API for 21 languages
  • 4 tiers · nano · mini · large · Speaker labels
नमस्ते, मेरा ऑर्डर कब आएगा?

Hindi · 8 kHz phone call → text, in real time

₹20

per hour of audio · $0.208

Pica

Speaks

Natural text to speech in Hindi and English, with expressive voices for assistants, IVR and narration.

  • text-to-speech API for Hindi and English
  • 3 tiers · nano · small · large

“Welcome back, Riya! Your order is on its way.”

₹21

per hour of audio · $0.219

Osprey Live

Talks

Real-time multimodal AI that sees, hears, thinks and talks — with tool calling, for AI calling, companions, robots, glasses and cars.

  • real-time multimodal AI model
Live call · 00:42

Caller: “Do you deliver on Sundays?”

Osprey Live: “Yes — until 9 pm. Shall I book you a slot?”

₹65.74

per call-hour, all in · $0.684

Embeddings & Reranking Finds

Semantic search and RAG — embed your documents, rerank a shortlist, and give Osprey your own knowledge to answer from.

Docs