CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alvanlii /cantonese-radiogated Cantonese Radio Pseudo-Transcription Dataset Contains 14k hours of audio sourced from Archive.org Columns order_index: Represents the order of the audio compared to those from the same filename link: Link of the original full audio transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-radio.audioautomatic-speech-recognition1M<n<10M25 likes2.3k downloads2y agoHugging Face02alvanlii /cantonese-youtube-ttsgated Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet Filtered out: Overlapped voices, detected using pyannote/speaker-diarization-3.1 Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.audiotext-to-speech1M<n<10M3 likes873 downloads6mo agoHugging Face03alvanlii /cantonese-youtubegated Cantonese Youtube Pseudo-Transcription Dataset Contains approximately 10k hours of audio sourced from YouTube Videos are chosen at random, and scraped on a channel basis Includes news, vlogs, entertainment, stories, health Columns transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube.audioautomatic-speech-recognition1M<n<10M51 likes607 downloads2y agoHugging Face04OrcinusOrca /YouTube-Cantonese Cantonese Audio Dataset from YouTube This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models. Data Source and Processing The data was obtained through the following process: Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.audioautomatic-speech-recognition100K<n<1M5 likes445 downloads1y agoHugging Face05Yvthyvq /KarenSo_CantoneseRecordings_Liujgoj KarenSo_CantoneseRecordings_Liujgoj 本數據集係基於開源粵語語音庫 kakiso/KarenSo_CantoneseRecordings 進行重構與正字、拼音重映射嘅溜歌粵語(Liujgoj)語音正字 Ground Truth 數據集。主要用於音標切分、粵語語音流形(Language Manifold)對齊、以及 Stage 2 SFT 翻譯與語言工程任務。 👥 致謝與上游數據集說明 (Acknowledgment & Upstream Source) 本數據集嘅原始音頻與文本來源於 Hugging Face 社群成員 kakiso 分享嘅項目: 原始數據集 (Original Dataset): kakiso/KarenSo_CantoneseRecordings 原始授權協議 (License): CC-BY-4.0 在此由衷感謝原創作者 Karen So 及其團隊錄製並無私分享高品質(44.1kHz / 16bit / Mono)嘅純淨粵語口語語料,為廣東話開源 AI… See the full description on the dataset page: https://huggingface.co/datasets/Yvthyvq/KarenSo_CantoneseRecordings_Liujgoj.audioautomatic-speech-recognition0 likes162 downloads3mo agoHugging Face06psdn-ai /cantonese-speech-samplesgated Cantonese Speech Samples This sample shows Cantonese speech with native transcript metadata. It is meant to help buyers review dialect fit, recording quality, and sample structure before requesting broader coverage. What This Shows Cantonese speech audio with paired transcript metadata Language-specific metadata for review and delivery planning A compact preview of the available sample structure Dataset Specifications Field Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/cantonese-speech-samples.audioautomatic-speech-recognitionn<1K0 likes15 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.