cantonese
cantonese-radio
Cantonese Radio Pseudo-Transcription Dataset
Contains 14k hours of audio sourced from Archive.org
Columns
order_index: Represents the order of the audio compared to those from the same filename
link: Link of the original full audio
transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding
transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall
used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-radio.cantonese_dailyCantonese_Common_Crawl_Filtered
Cantonese Chinese C4
Dataset Summary
Downloaded and processed using code based on another project attempting to recreate the C4 dataset.
The resultant traditional Chinese dataset can be found here.
This dataset contains data processed with CantoneseDetect.
In CantoneseDetect, you can choose whether to include quotes (i.e. categorise the data as Cantonese even if Cantonese appeared only in quotes).
And I found that a lot of entries came from Wikipedia and LIHKG. If you… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese_Common_Crawl_Filtered.cantonese-youtube-tts
Cantonese Audio TTS Dataset
This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet
Filtered out:
Overlapped voices, detected using pyannote/speaker-diarization-3.1
Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.YouTube-Cantonese-Emilia
YouTube Cantonese — Emilia
2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by
running alvanlii/cantonese-youtube
through the Emilia
speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering).
Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn
label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in
both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.sbs_cantonese
SBS Cantonese Speech Corpus
This speech corpus contains 435 hours of SBS Cantonese podcasts from Auguest 2022 to October 2023.
There are 2,519 episodes and each episode is split into segments that are at most 10 seconds long. In total, there are 189,216 segments in this corpus.
Here is a breakdown on the categories of episodes present in this dataset:
Category
SBS Channels
Episodes
news
中文新聞, 新聞簡報
622
business
寰宇金融
148
vaccine
疫苗快報
71
gardening
園藝趣談
58
tech
科技世界… See the full description on the dataset page: https://huggingface.co/datasets/AlienKevin/sbs_cantonese.
