Subanana
Cantonese speech to text · 廣東話

Cantonese audio to text.

Upload a Cantonese recording — MP3, M4A, WAV, a video file, or a public YouTube link — and get it back as text: speakers separated, punctuation restored, every line timecoded. Choose colloquial Cantonese (口語) or Standard Written Chinese (書面語) as the output. Subanana averages 98% accuracy across the 95+ languages it supports. Preview the first 15 minutes of any file free.

Start free15-minute free preview · no credit card
A still lake at first light
plus.subanana.com
Transcript
Speaker 100:09:12

你哋喺灣仔開第一間舖,當初係咩原因?

Speaker 200:09:18

主要係人流穩定,加上租金我哋負擔得起。

Speaker 100:09:27

頭三個月最難嘅係咩?

Speaker 200:09:33

請人最難,全職請咗成兩個月先請到。

  • Google
  • Deloitte
  • dentsu
  • Manulife
  • NAVER
  • Philips
  • Amazon
  • Shopify
  • Figma
  • Coinbase
  • WPP
  • Semrush
  • Google
  • Deloitte
  • dentsu
  • Manulife
  • NAVER
  • Philips
  • Amazon
  • Shopify
  • Figma
  • Coinbase
  • WPP
  • Semrush

How does it work?

interview-ep12.m4a

M4A · 27:45 · Uploaded

MP3, M4A or WAV audio, MP4 or MOV video, or a public linkUp to 8 hours and 30 GB per file, on every plan

Transcribe Cantonese

00:14

今日想同大家傾下新分店嘅籌備進度。

00:22

裝修下星期開始,預計十月頭完工。

00:30

招聘方面,全職請咗四個,仲爭兩個。

SpeakersSpeaker 1Speaker 2Punctuation restored

Text formats

TXTDOCXXLSXMarkdown

Cantonese transcript

TXT · DOCX · XLSX · Markdown

Written Chinese (書面語)

from spoken Cantonese

Subtitle files

SRT · VTT

English translation

and 95+ other languages

Ask the recording

answers cite their cue

Nine reasons to stop transcribing Cantonese recordings by hand.

From an audio file or a link, through recognition and proofreading, to the text file you actually need — all of it in one place.

Bring the recording in from anywhere

  • Upload the fileMP3, M4A, WAV — or a video: MP4, MOV, WEBM. Drag it into the browser.
  • Or paste a public linkYouTube, Instagram and Facebook links work without downloading the video first.
  • Long recordings stay wholeUp to 8 hours and 30 GB per file, on every plan.
  • Record in the browserCapture a call or an interview right here, then convert it.

Cantonese is a supported language in its own right

  • Pick Cantonese as the source廣東話 is its own entry in the language picker — the audio is not treated as Mandarin and converted afterwards.
  • Spoken or written outputChoose colloquial Cantonese (口語) or Standard Written Chinese (書面語), each in Traditional or Simplified characters.
  • Cantonese mixed with EnglishEnglish words inside a Cantonese sentence stay as English, spelled properly.
  • Translate while you are hereThe same recording can also come back as text in English or 95+ other languages.

Whatever text file you need, it exports

  • Text and tablesTXT, DOCX, XLSX and Markdown — speakers and timestamps included.
  • Subtitle filesSRT and VTT, single-language or bilingual.
  • All at onceDownload every format together as a ZIP.
  • Share a linkSend the result to someone without making them open an account.

Invented sentences do not reach your transcript

  • Two layers of detectionAI hallucination detection plus an engine-side check runs over every segment; suspect lines get flagged.
  • Automatic re-run on another engineFlagged segments go to a different engine instead of being written through.
  • Silence stays silenceA pause is a pause, not a sentence the model imagined.

Proofread it before anyone reads it

  • A score out of 100Suggestions grouped by kind, usually back in ten to thirty seconds.
  • Approve group by groupOnly the groups you accept get applied; the rest stay untouched.
  • Timecodes do not moveProofreading changes wording, never the timing.
  • Entirely skippableNothing changes until you press the button.

Names and jargon come out spelled right

  • Workspace glossarySave people, products and in-house terms once; every later file uses your spelling.
  • Per project or everywhereMark a term universal, or attach it to a single project.
  • Import what you already haveLoad a CSV or XLSX, or let the AI pull candidates from a PDF.
  • Review before it appliesSee exactly which terms were aligned, before anything changes.

The same transcript gives you more than a transcript

  • Summaries and highlights30+ templates pull the points out of a long recording.
  • Decisions and action itemsMeeting recordings come back with what was decided and who took it.
  • Descriptions and notesVideo descriptions and bullet lists come from the same text.
  • Ask the recordingAsk the AI assistant directly — every answer points back to its cue.

Need subtitles too? Style them here

  • Two style basesOutline with a shadow, or a background block with adjustable opacity.
  • Adaptive sizingLong lines reflow instead of running off the frame.
  • Two tracks, styled apartSource and translation never fight for the same line.
  • Preview as you goFont, size, colour, alignment and padding, on the real frame.

Want the finished video, not just the text?

  • Up to 4K720P, 1080P, 2K or 4K, rendered exactly as previewed.
  • Bilingual burn-inCantonese and English on screen together, each with its own styling.
  • Your watermark, or noneHide the default watermark entirely, or swap in your own.
  • Renders in the backgroundKeep editing while it runs; the export is the version you pressed on.

If the recording is in Cantonese and you need it as text, this is for you.

Four kinds of audio, one problem — typing out an hour of Cantonese speech by hand takes three.

Interviews and press conferences

Journalists and editors

Quotes have to be exact and attributed, with the minute of the recording they came from.

Fieldwork, oral history and lectures

Researchers and students

Cantonese interviews sit unquotable until someone types them out — and nobody has three hours per tape to re-listen.

Meetings and client calls

Teams and businesses

Get a searchable text record of what was said — as spoken Cantonese, or as Standard Written Chinese for the minutes.

YouTube channels and podcasts

Creators and podcasters

Turn the episode into text once, then reuse it for the description, the blog version and the captions.

Everything else.

None of these is why you would switch, but every one of them is there when you have to hand something over.

File uploadPaste a linkBrowser recordingScreen recordingSpeaker detectionCustom glossaryFind and replaceWaveform timelineAI proofreadAI summariesChapter markersAsk the videoSRT · VTT exportDOCX · XLSX exportTXT · Markdown exportDownload everything as a ZIPShare links without sign-inWorkspace roles

The details live in the product — a free workspace takes about a minute.

See it in the product

Before you startthe questions everyone asks

Yes. Cantonese (廣東話) is a supported source language in its own right — you pick it as the language of the recording, and the system routes the audio to the model that benchmarks best for it. It is not transcribed as Mandarin and converted afterwards.

Your choice, at the language step. Colloquial Cantonese (口語, e.g. 佢哋而家搞掂咗) keeps what was said the way it was said; Standard Written Chinese (書面語, e.g. 他們現在處理好了) reads like formal writing. Each comes in Traditional or Simplified characters — four combinations in total.

A free account uploads a recording and previews the first 15 minutes of the text for each file, 3 files a month, with no credit card. Downloading the transcript or subtitle file is a paid feature — so you can check the output on your own Cantonese audio before paying anything.

MP3, M4A and WAV recordings, plus video files like MP4, MOV and WEBM — a phone voice memo or a meeting recording uploads as-is, no conversion first. Each file can run up to 8 hours and 30 GB, on every plan including free.

Yes. Paste a public YouTube, Instagram or Facebook link and Subanana fetches and transcribes it, so you do not have to download the video first. Note that Subanana returns the transcript — it does not hand you the third-party video file.

Hong Kong speech rarely stays in one language, and mixed sentences are handled as such: English product names, company names and phrases inside a Cantonese sentence come out as English, spelled properly, not as phonetic guesses.

Subanana averages 98% accuracy. What you get on any one recording moves with clarity, background noise and jargon, and adding names to the workspace glossary visibly improves how people and companies come out.

Every segment passes AI hallucination detection and an engine-side check; anything suspect is flagged and re-run on a different engine rather than written through. A pause stays a pause. You can also run an AI proofread before exporting and approve the suggestions group by group.

Yes. Once the transcript exists, English is one of 95+ translation targets, and a bilingual subtitle file can carry the Cantonese line and the English line together.