datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cantonese-radio
Cantonese Radio Pseudo-Transcription Dataset
Contains 14k hours of audio sourced from Archive.org
Columns
order_index: Represents the order of the audio compared to those from the same filename
link: Link of the original full audio
transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding
transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall
used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-radio.cantonese-youtube-tts
Cantonese Audio TTS Dataset
This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet
Filtered out:
Overlapped voices, detected using pyannote/speaker-diarization-3.1
Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.cantonese-youtube
Cantonese Youtube Pseudo-Transcription Dataset
Contains approximately 10k hours of audio sourced from YouTube
Videos are chosen at random, and scraped on a channel basis
Includes news, vlogs, entertainment, stories, health
Columns
transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding
transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall
used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube.YouTube-Cantonese
Cantonese Audio Dataset from YouTube
This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models.
Data Source and Processing
The data was obtained through the following process:
Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.KarenSo_CantoneseRecordings_Liujgoj
KarenSo_CantoneseRecordings_Liujgoj
本數據集係基於開源粵語語音庫 kakiso/KarenSo_CantoneseRecordings 進行重構與正字、拼音重映射嘅溜歌粵語(Liujgoj)語音正字 Ground Truth 數據集。主要用於音標切分、粵語語音流形(Language Manifold)對齊、以及 Stage 2 SFT 翻譯與語言工程任務。
👥 致謝與上游數據集說明 (Acknowledgment & Upstream Source)
本數據集嘅原始音頻與文本來源於 Hugging Face 社群成員 kakiso 分享嘅項目:
原始數據集 (Original Dataset): kakiso/KarenSo_CantoneseRecordings
原始授權協議 (License): CC-BY-4.0
在此由衷感謝原創作者 Karen So 及其團隊錄製並無私分享高品質(44.1kHz / 16bit / Mono)嘅純淨粵語口語語料,為廣東話開源 AI… See the full description on the dataset page: https://huggingface.co/datasets/Yvthyvq/KarenSo_CantoneseRecordings_Liujgoj.cantonese-speech-samples
Cantonese Speech Samples
This sample shows Cantonese speech with native transcript metadata. It is meant to help buyers review dialect fit, recording quality, and sample structure before requesting broader coverage.
What This Shows
Cantonese speech audio with paired transcript metadata
Language-specific metadata for review and delivery planning
A compact preview of the available sample structure
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/cantonese-speech-samples.
