datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sbs_cantonese
SBS Cantonese Speech Corpus
This speech corpus contains 435 hours of SBS Cantonese podcasts from Auguest 2022 to October 2023.
There are 2,519 episodes and each episode is split into segments that are at most 10 seconds long. In total, there are 189,216 segments in this corpus.
Here is a breakdown on the categories of episodes present in this dataset:
Category
SBS Channels
Episodes
news
中文新聞, 新聞簡報
622
business
寰宇金融
148
vaccine
疫苗快報
71
gardening
園藝趣談
58
tech
科技世界… See the full description on the dataset page: https://huggingface.co/datasets/AlienKevin/sbs_cantonese.YouTube-Cantonese
Cantonese Audio Dataset from YouTube
This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models.
Data Source and Processing
The data was obtained through the following process:
Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.
