Subanana
Chinese audio to text · Mandarin & more

Chinese audio to text.

Upload a Chinese recording — MP3, M4A, WAV or a video — and this converter turns the speech into text: a real transcript with speakers separated, punctuation restored and every line timecoded. Mandarin, Cantonese and the rest of 95+ languages, at 98% average accuracy. Preview the first 15 minutes of any file free.

Start free15-minute free preview · no credit card
A coastline in soft morning light
plus.subanana.com
Transcript
Speaker 100:08:12

这个 campaign 下周三就要上线。

Speaker 200:08:18

设计那边还差两个 banner,明天补齐。

Speaker 100:08:27

预算方面需要调整吗?

Speaker 200:08:33

维持之前定好的数字,不用加。

  • Google
  • Deloitte
  • dentsu
  • Manulife
  • NAVER
  • Philips
  • Amazon
  • Shopify
  • Figma
  • Coinbase
  • WPP
  • Semrush
  • Google
  • Deloitte
  • dentsu
  • Manulife
  • NAVER
  • Philips
  • Amazon
  • Shopify
  • Figma
  • Coinbase
  • WPP
  • Semrush

How does it work?

team-sync.m4a

M4A · 38:45 · uploaded

Upload a recording, or paste a public linkMandarin and Cantonese are separate languages here

Chinese transcript

00:12

这个 campaign 下周三就要上线。

00:18

设计那边还差两个 banner,明天补齐。

00:27

预算维持之前定好的数字。

Speakers separatedSpeaker 1Speaker 2timecoded

Output options

TraditionalSimplifiedTXT · DOCXXLSX · SRT

Chinese transcript

TXT · DOCX · XLSX · Markdown

Traditional or Simplified

characters, your choice

Subtitle files

SRT · VTT

English translation

and 95+ other languages

Summary and answers

key points · ask the audio

Nine reasons to stop transcribing Chinese audio by hand.

From a recording or a link, through recognition and proofreading, to the text file you actually need — all of it in one place.

Bring the video in from anywhere

  • Upload the fileMP4, MOV, WEBM, MP3, M4A, WAV — drag it into the browser.
  • Or paste a public linkYouTube, Instagram and Facebook links work without downloading the video first.
  • Long recordings stay wholeUp to 8 hours and 30 GB per file, on every plan.
  • Record in the browserCapture a call or a screen right here, then convert it.

It knows the language your video is in

  • 95+ languagesPick the source language and the system routes to the model that benchmarks best for it.
  • Marathi includedमराठी is one of the supported source languages, not a machine-translated afterthought.
  • Mixed-language speechEnglish words inside a Marathi sentence stay as English, spelled properly.
  • Translate while you are hereThe same recording can also come back as text in another language.

Whatever text file you need, it exports

  • Text and tablesTXT, DOCX, XLSX and Markdown — speakers and timestamps included.
  • Subtitle filesSRT and VTT, single-language or bilingual.
  • All at onceDownload every format together as a ZIP.
  • Share a linkSend the result to someone without making them open an account.

Invented sentences do not reach your transcript

  • Two layers of detectionAI hallucination detection plus an engine-side check runs over every segment; suspect lines get flagged.
  • Automatic re-run on another engineFlagged segments go to a different engine instead of being written through.
  • Silence stays silenceA pause is a pause, not a sentence the model imagined.

Proofread it before anyone reads it

  • A score out of 100Suggestions grouped by kind, usually back in ten to thirty seconds.
  • Approve group by groupOnly the groups you accept get applied; the rest stay untouched.
  • Timecodes do not moveProofreading changes wording, never the timing.
  • Entirely skippableNothing changes until you press the button.

Names and jargon come out spelled right

  • Workspace glossarySave people, products and in-house terms once; every later file uses your spelling.
  • Per project or everywhereMark a term universal, or attach it to a single project.
  • Import what you already haveLoad a CSV or XLSX, or let the AI pull candidates from a PDF.
  • Review before it appliesSee exactly which terms were aligned, before anything changes.

The same transcript gives you more than a transcript

  • Summaries and highlights30+ templates pull the points out of a long recording.
  • Chapters and quotesLong videos get split automatically; quotable lines are ready to use.
  • Descriptions and notesVideo descriptions and bullet lists come from the same text.
  • Ask the videoAsk the AI assistant directly — every answer points back to its cue.

Need subtitles too? Style them here

  • Two style basesOutline with a shadow, or a background block with adjustable opacity.
  • Adaptive sizingLong lines reflow instead of running off the frame.
  • Two tracks, styled apartSource and translation never fight for the same line.
  • Preview as you goFont, size, colour, alignment and padding, on the real frame.

Want the finished video, not just the text?

  • Up to 4K720P, 1080P, 2K or 4K, rendered exactly as previewed.
  • Bilingual burn-inBoth languages on screen, each with its own styling.
  • Your watermark, or noneHide the default watermark entirely, or swap in your own.
  • Renders in the backgroundKeep editing while it runs; the export is the version you pressed on.

If the recording is in Chinese and you need it in text, this is for you.

Four kinds of audio, one problem — an hour of speech takes three to type out, in any language.

Meetings in Mandarin or Cantonese

Teams working across languages

The meeting happened in Chinese; the record needs to work for everyone. Transcribe it, then read or translate it in the language you work in.

Interviews and fieldwork

Students and researchers

Chinese-language interviews become searchable, quotable text — with the exact minute each quote came from.

Pressers and recorded calls

Journalists and editors

Quotes must be exact and attributed. Speakers are separated and every line is timecoded.

Chinese-language shows

Creators and podcasters

One transcription feeds the captions, the show notes and the translated version for a second audience.

Everything else.

None of these is why you would switch, but every one of them is there when you have to hand something over.

File uploadPaste a linkBrowser recordingScreen recordingSpeaker detectionCustom glossaryFind and replaceWaveform timelineAI proofreadAI summariesChapter markersAsk the videoSRT · VTT exportDOCX · XLSX exportTXT · Markdown exportDownload everything as a ZIPShare links without sign-inWorkspace roles

The details live in the product — a free workspace takes about a minute.

See it in the product

Before you startthe questions everyone asks

Mandarin is the default for Chinese audio, and Cantonese is supported as a separate language in its own right rather than being transcribed as Mandarin first. If your recording is Cantonese, the dedicated Cantonese audio to text page covers that job in more detail, including the spoken-style and written-Chinese output choice.

A free account uploads a file and previews the first 15 minutes of the text, 3 files a month, no credit card. Downloading the transcript or subtitle file is a paid feature — so you can check the output on your own recording before paying anything.

Yes. Code-switched speech — an English term dropped mid-sentence — is recognised in place and kept in the original English, in one clean line rather than two broken ones.

Yes. Once the Chinese transcript exists, English is one of 95+ translation targets, and a bilingual export can carry the Chinese line and the English line together.

Three steps: upload the recording (or paste a public link), pick the spoken language, and the AI transcribes it. Review the transcript in the editor, then export it as TXT, DOCX, XLSX or SRT.

MP3, M4A and WAV, plus video files like MP4, MOV and WEBM — video is transcribed the same way. Each file can run up to 8 hours and 30 GB, on every plan. You can also paste a public YouTube, Instagram or Facebook link.

98% on average. Clarity, background noise and jargon all matter, and registering names in a glossary noticeably steadies proper nouns.

Transcripts export as TXT, DOCX, XLSX or Markdown; subtitles as SRT or VTT — or every format at once in a ZIP.

Every segment passes AI hallucination detection and an engine-side check; anything suspect is flagged and re-run on a different engine rather than written through. A pause stays a pause. You can also run an AI proofread and approve its suggestions group by group before exporting.