Meetings in Mandarin or Cantonese
Teams working across languages
The meeting happened in Chinese; the record needs to work for everyone. Transcribe it, then read or translate it in the language you work in.
Upload a Chinese recording — MP3, M4A, WAV or a video — and this converter turns the speech into text: a real transcript with speakers separated, punctuation restored and every line timecoded. Mandarin, Cantonese and the rest of 95+ languages, at 98% average accuracy. Preview the first 15 minutes of any file free.

这个 campaign 下周三就要上线。
设计那边还差两个 banner,明天补齐。
预算方面需要调整吗?
维持之前定好的数字,不用加。
























team-sync.m4a
M4A · 38:45 · uploaded
Chinese transcript
这个 campaign 下周三就要上线。
设计那边还差两个 banner,明天补齐。
预算维持之前定好的数字。
Output options
TraditionalSimplifiedTXT · DOCXXLSX · SRTChinese transcript
TXT · DOCX · XLSX · Markdown
Traditional or Simplified
characters, your choice
Subtitle files
SRT · VTT
English translation
and 95+ other languages
Summary and answers
key points · ask the audio
From a recording or a link, through recognition and proofreading, to the text file you actually need — all of it in one place.
Four kinds of audio, one problem — an hour of speech takes three to type out, in any language.
Meetings in Mandarin or Cantonese
The meeting happened in Chinese; the record needs to work for everyone. Transcribe it, then read or translate it in the language you work in.
Interviews and fieldwork
Chinese-language interviews become searchable, quotable text — with the exact minute each quote came from.
Pressers and recorded calls
Quotes must be exact and attributed. Speakers are separated and every line is timecoded.
Chinese-language shows
One transcription feeds the captions, the show notes and the translated version for a second audience.
None of these is why you would switch, but every one of them is there when you have to hand something over.
The details live in the product — a free workspace takes about a minute.
See it in the productMandarin is the default for Chinese audio, and Cantonese is supported as a separate language in its own right rather than being transcribed as Mandarin first. If your recording is Cantonese, the dedicated Cantonese audio to text page covers that job in more detail, including the spoken-style and written-Chinese output choice.
A free account uploads a file and previews the first 15 minutes of the text, 3 files a month, no credit card. Downloading the transcript or subtitle file is a paid feature — so you can check the output on your own recording before paying anything.
Yes. Code-switched speech — an English term dropped mid-sentence — is recognised in place and kept in the original English, in one clean line rather than two broken ones.
Yes. Once the Chinese transcript exists, English is one of 95+ translation targets, and a bilingual export can carry the Chinese line and the English line together.
Three steps: upload the recording (or paste a public link), pick the spoken language, and the AI transcribes it. Review the transcript in the editor, then export it as TXT, DOCX, XLSX or SRT.
MP3, M4A and WAV, plus video files like MP4, MOV and WEBM — video is transcribed the same way. Each file can run up to 8 hours and 30 GB, on every plan. You can also paste a public YouTube, Instagram or Facebook link.
98% on average. Clarity, background noise and jargon all matter, and registering names in a glossary noticeably steadies proper nouns.
Transcripts export as TXT, DOCX, XLSX or Markdown; subtitles as SRT or VTT — or every format at once in a ZIP.
Every segment passes AI hallucination detection and an engine-side check; anything suspect is flagged and re-run on a different engine rather than written through. A pause stays a pause. You can also run an AI proofread and approve its suggestions group by group before exporting.
Mandarin is the default for Chinese audio, and Cantonese is supported as a separate language in its own right rather than being transcribed as Mandarin first. If your recording is Cantonese, the dedicated Cantonese audio to text page covers that job in more detail, including the spoken-style and written-Chinese output choice.
A free account uploads a file and previews the first 15 minutes of the text, 3 files a month, no credit card. Downloading the transcript or subtitle file is a paid feature — so you can check the output on your own recording before paying anything.
Yes. Code-switched speech — an English term dropped mid-sentence — is recognised in place and kept in the original English, in one clean line rather than two broken ones.
Yes. Once the Chinese transcript exists, English is one of 95+ translation targets, and a bilingual export can carry the Chinese line and the English line together.
Three steps: upload the recording (or paste a public link), pick the spoken language, and the AI transcribes it. Review the transcript in the editor, then export it as TXT, DOCX, XLSX or SRT.
MP3, M4A and WAV, plus video files like MP4, MOV and WEBM — video is transcribed the same way. Each file can run up to 8 hours and 30 GB, on every plan. You can also paste a public YouTube, Instagram or Facebook link.
98% on average. Clarity, background noise and jargon all matter, and registering names in a glossary noticeably steadies proper nouns.
Transcripts export as TXT, DOCX, XLSX or Markdown; subtitles as SRT or VTT — or every format at once in a ZIP.
Every segment passes AI hallucination detection and an engine-side check; anything suspect is flagged and re-run on a different engine rather than written through. A pause stays a pause. You can also run an AI proofread and approve its suggestions group by group before exporting.
Stop retyping what was said.