Want the output, not the API
Creators who read the launch news
The model made headlines; your deliverable is still a subtitled video. Upload, review the cues, export — the engine layer is already handled.
Google's new model made speech-to-text headline news — as a developer API. If what you want is a subtitled video, upload it here instead: frontier-class engines benchmarked per language, cues timed and checked, SRT or burned-in captions out in minutes. No API key, no code. Free 15-minute preview per file.
No card needed. First 15 minutes free.
























Launch-day review
MP4 · 14:02 · uploaded
Timed subtitle cues
Let’s talk about the new transcription models.
Accuracy jumped again this year.
But most of them are APIs, not apps.
Download
launch-review-en.srtlaunch-review-fr.srtall-formats.zipSubtitle file
SRT · VTT
Bilingual subtitles
source + translation, cue by cue
Burned-in video
captions on the frame · up to 4K
Translation
into 95+ languages
Text files
TXT · DOCX · XLSX · Markdown
Every row below is a piece of the pipeline you would otherwise build yourself.
If the deliverable is a captioned video, the API was never the requirement.
Want the output, not the API
The model made headlines; your deliverable is still a subtitled video. Upload, review the cues, export — the engine layer is already handled.
Prototype before you build
See what production-grade cue segmentation, reading-speed flags and bilingual exports look like before committing to writing your own.
Hours of recordings
Long files process in one pass — up to 8 hours and 30 GB each — and come back captioned, searchable and ready to publish.
More than one language
Transcribe the source, translate into any of 95+ targets, and ship bilingual subtitles or burned-in captions per market.
None of these are the reason to switch, but every one of them is there when a delivery spec asks for it.
The detail lives in the product — a free workspace takes about a minute to set up.
See it in the productSubanana doesn't lock to any single vendor. Every file is routed across continuously benchmarked frontier speech engines — the same class as the models in the headlines — picking the best performer for its source language, and hallucination checks re-run suspect segments on a different engine. As new models ship, the benchmark decides, not the branding.
Per Google's launch post, the model returns transcription data — text with word-level timestamps and speaker attribution — through the Gemini API. Turning that into an SRT file means writing cue-segmentation and formatting code. Here, it's the export button.
Get an API key in Google AI Studio, send pre-recorded audio to the model through the Interactions API — it auto-detects 85+ languages and attributes up to 3 speakers, per Google — then convert the returned words and timestamps into subtitle cues in your own code. The right path if you're building transcription into a product of your own.
Per the launch post: automatic detection of 85+ languages including regional accents, word-level timestamps with speaker attribution for pre-recorded audio, support for up to 3 speakers with more marked experimental, real-time streaming through the Live API, and a Google-reported average word error rate of 2.6% for pre-recorded audio and 4.0% for streaming.
Overall accuracy reaches 98%. Every chunk passes two hallucination checks — silence stays silent instead of becoming an invented sentence — and an optional proofread pass proposes misheard and same-sounding words for you to approve, without moving a single timestamp.
SRT, VTT, TXT, DOCX, XLSX and Markdown — single-language or bilingual, or one ZIP with everything. You can also export the video itself with the captions burned in, up to 4K.
A free account previews the first 15 minutes of each file, 3 files a month, no credit card. Downloading the subtitle file is a paid feature — you check the quality on your own video before paying anything.
Subanana doesn't lock to any single vendor. Every file is routed across continuously benchmarked frontier speech engines — the same class as the models in the headlines — picking the best performer for its source language, and hallucination checks re-run suspect segments on a different engine. As new models ship, the benchmark decides, not the branding.
Per Google's launch post, the model returns transcription data — text with word-level timestamps and speaker attribution — through the Gemini API. Turning that into an SRT file means writing cue-segmentation and formatting code. Here, it's the export button.
Get an API key in Google AI Studio, send pre-recorded audio to the model through the Interactions API — it auto-detects 85+ languages and attributes up to 3 speakers, per Google — then convert the returned words and timestamps into subtitle cues in your own code. The right path if you're building transcription into a product of your own.
Per the launch post: automatic detection of 85+ languages including regional accents, word-level timestamps with speaker attribution for pre-recorded audio, support for up to 3 speakers with more marked experimental, real-time streaming through the Live API, and a Google-reported average word error rate of 2.6% for pre-recorded audio and 4.0% for streaming.
Overall accuracy reaches 98%. Every chunk passes two hallucination checks — silence stays silent instead of becoming an invented sentence — and an optional proofread pass proposes misheard and same-sounding words for you to approve, without moving a single timestamp.
SRT, VTT, TXT, DOCX, XLSX and Markdown — single-language or bilingual, or one ZIP with everything. You can also export the video itself with the captions burned in, up to 4K.
A free account previews the first 15 minutes of each file, 3 files a month, no credit card. Downloading the subtitle file is a paid feature — you check the quality on your own video before paying anything.
Stop retyping what was said.