CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alvanlii /cantonese-radiogated Cantonese Radio Pseudo-Transcription Dataset Contains 14k hours of audio sourced from Archive.org Columns order_index: Represents the order of the audio compared to those from the same filename link: Link of the original full audio transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-radio.audioautomatic-speech-recognition1M<n<10M25 likes2.2k downloads2y agoHugging Face02ziyou-li /cantonese_dailyaudio1K<n<10K4 likes1.5k downloads4y agoHugging Face03jed351 /Cantonese_Common_Crawl_Filtered Cantonese Chinese C4 Dataset Summary Downloaded and processed using code based on another project attempting to recreate the C4 dataset. The resultant traditional Chinese dataset can be found here. This dataset contains data processed with CantoneseDetect. In CantoneseDetect, you can choose whether to include quotes (i.e. categorise the data as Cantonese even if Cantonese appeared only in quotes). And I found that a lot of entries came from Wikipedia and LIHKG. If you… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese_Common_Crawl_Filtered.text1M<n<10M4 likes1.1k downloads1y agoHugging Face04alvanlii /cantonese-youtube-ttsgated Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet Filtered out: Overlapped voices, detected using pyannote/speaker-diarization-3.1 Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.audiotext-to-speech1M<n<10M3 likes804 downloads6mo agoHugging Face05Scicom-intl /YouTube-Cantonese-Emilia YouTube Cantonese — Emilia 2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by running alvanlii/cantonese-youtube through the Emilia speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering). Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.tabularautomatic-speech-recognition1M<n<10M1 likes790 downloads1mo agoHugging Face06AlienKevin /sbs_cantonese SBS Cantonese Speech Corpus This speech corpus contains 435 hours of SBS Cantonese podcasts from Auguest 2022 to October 2023. There are 2,519 episodes and each episode is split into segments that are at most 10 seconds long. In total, there are 189,216 segments in this corpus. Here is a breakdown on the categories of episodes present in this dataset: Category SBS Channels Episodes news 中文新聞, 新聞簡報 622 business 寰宇金融 148 vaccine 疫苗快報 71 gardening 園藝趣談 58 tech 科技世界… See the full description on the dataset page: https://huggingface.co/datasets/AlienKevin/sbs_cantonese.audio10K<n<100K7 likes789 downloads3y agoHugging Face07alvanlii /cantonese-youtubegated Cantonese Youtube Pseudo-Transcription Dataset Contains approximately 10k hours of audio sourced from YouTube Videos are chosen at random, and scraped on a channel basis Includes news, vlogs, entertainment, stories, health Columns transcript_whisper: Transcribed using Scrya/whisper-large-v2-cantonese with alvanlii/whisper-small-cantonese for speculative decoding transcript_sensevoice: Transcribed using FunAudioLLM/SenseVoiceSmall used OpenCC to convert to traditional chinese… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube.audioautomatic-speech-recognition1M<n<10M51 likes604 downloads2y agoHugging Face08OrcinusOrca /YouTube-Cantonese Cantonese Audio Dataset from YouTube This dataset contains Cantonese audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models. Data Source and Processing The data was obtained through the following process: Download: Audio (.m4a) and available Cantonese subtitles (.srt for zh-TW, zh-HK, zh-Hant)… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-Cantonese.audioautomatic-speech-recognition100K<n<1M5 likes439 downloads1y agoHugging Face09kennychan-6 /Cantonese_Datasetaudio1K<n<10K0 likes404 downloads1y agoHugging Face10ziyou-li /cantonese_processed Dataset Card for "cantonese_processed" More Information needed 10K<n<100K0 likes360 downloads4y agoHugging Face11AlienKevin /mixed_cantonese_and_english_speechThe Mixed Cantonese and English (MCE) dataset covers 18 topics related to daily life, comprising a total of 34.8 hours of audio files. The corresponding annotated text consists of 307,540 Chinese characters and 70,132 English words. Among the topics, the "Food" category has the highest frequency of English words, with a Chinese character to English word ratio of approximately 3:1. On the other hand, the "Tech News" topic has the lowest frequency of English words, approximately 8:1. We randomly… See the full description on the dataset page: https://huggingface.co/datasets/AlienKevin/mixed_cantonese_and_english_speech.audio10K<n<100K19 likes265 downloads2y agoHugging Face12NiuChou /pazhou-cantonese-asr-16kaudio1K<n<10K0 likes244 downloads2mo agoHugging Face13MagicHub /magicdata-dialect-cantonese-tts-lite MagicData-Dialect-Northeastern Chinese-TTS-Lite MAGIC DATA OPEN-SOURCE LICENSE Dataset Overview Item Information Dataset Type N/A Language Chinese Dialect Speech Style Scripted Content N/A Audio Parameters 48 kHz, 16 bits File Format WAV (PCM) Recording Equipment microphone Recording Environment quiet indoor environment License Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/magicdata-dialect-cantonese-tts-lite.audion<1K1 likes213 downloads3mo agoHugging Face14raptorkwok /cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences. texttranslation100K<n<1M17 likes194 downloads3y agoHugging Face15leeduckgo /cantonese-life-scenarios-corpus 多用途粤语生活场景语料库 1 项目简介 「多用途粤语生活场景语料库」是一个面向粤语学习、语音识别、语音合成、智能客服、对话系统和大模型训练的生活场景类有声语料库。 本语料库以真实日常生活中的高频表达为基础,围绕问候、感谢、做客、出行、住宿、就餐、购物、逛街、通讯、邮寄、银行、看病、天气、时间、旅游、运动、兴趣爱好、休闲娱乐、求职、工作、校园、租房、美容美发、问路、聚会应酬、约定约会、情绪表达、常用语、聊天、寻求帮助等常见场景,构建约 10,000 句粤语生活表达数据。 每条语料包含普通话文本、粤语文本、粤拼注音及粤语音频,适用于多种粤语人工智能应用场景。 2 目录结构 . ├── README.md # 项目说明文档(数据介绍、字段设计与使用方式) ├── index.csv # 语料索引表(句子级字段与场景信息的汇总) ├── template_pre.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/leeduckgo/cantonese-life-scenarios-corpus.audio1K<n<10K2 likes189 downloads3mo agoHugging Face16noncegeek /cantonese-drama-voice-fine-grained语料集名称:面向影视剧AI配音的粤语语料库 语料来源:AI DimSum Lab 简介: 本语料库是专为粤语影视剧 AI 配音模型训练构建的专用语料资源,核心涵盖《神雕侠侣(1995 古天乐版)》《乘龙怪婿》《寻秦记》等经典粤语影视内容。语料库匹配 AI 配音模型的人物识别、语音情绪识别、语音生成三大模块需求,并根据下游任务对《神雕侠侣》等语音数据进行了多情感、多人物、文本标注。其数据规模总计约800MB,影视时长超 7 小时,可提供丰富的粤语语音、情感、人物关联样本,能有效支撑模型训练中人物区分、情感还原、语音生成的精度提升,是粤语影视剧 AI 配音落地的核心数据基础。 适用场景: 粤语影视剧 AI 配音模型训练:直接用于模型的人物区分、情感还原、语音生成模块优化; 粤语语音研究:可作为粤语语音特征、情感语音分析的基础数据集; 影视 AI 技术开发:为影视领域的语音合成、角色语音克隆等技术提供数据支持。 使用说明: 本语料库仅用于非商业研究与技术开发(商业使用需联系维护者确认授权); 使用前建议对语音数据进行预处理(如降噪、采样率统一),以提升模型训练效果。 audio1K<n<10K2 likes187 downloads9mo agoHugging Face17AlienKevin /wordshk_cantonese_speechaudio100K<n<1M0 likes169 downloads2y agoHugging Face18Yvthyvq /KarenSo_CantoneseRecordings_Liujgoj KarenSo_CantoneseRecordings_Liujgoj 本數據集係基於開源粵語語音庫 kakiso/KarenSo_CantoneseRecordings 進行重構與正字、拼音重映射嘅溜歌粵語(Liujgoj)語音正字 Ground Truth 數據集。主要用於音標切分、粵語語音流形(Language Manifold)對齊、以及 Stage 2 SFT 翻譯與語言工程任務。 👥 致謝與上游數據集說明 (Acknowledgment & Upstream Source) 本數據集嘅原始音頻與文本來源於 Hugging Face 社群成員 kakiso 分享嘅項目: 原始數據集 (Original Dataset): kakiso/KarenSo_CantoneseRecordings 原始授權協議 (License): CC-BY-4.0 在此由衷感謝原創作者 Karen So 及其團隊錄製並無私分享高品質(44.1kHz / 16bit / Mono)嘅純淨粵語口語語料,為廣東話開源 AI… See the full description on the dataset page: https://huggingface.co/datasets/Yvthyvq/KarenSo_CantoneseRecordings_Liujgoj.audioautomatic-speech-recognition0 likes159 downloads3mo agoHugging Face19botisan-ai /cantonese-mandarin-translations Dataset Card for cantonese-mandarin-translations Dataset Summary This is a machine-translated parallel corpus between Cantonese (a Chinese dialect that is mainly spoken by Guangdong (province of China), Hong Kong, Macau and part of Malaysia) and Chinese (written form, in Simplified Chinese). Supported Tasks and Leaderboards N/A Languages Cantonese (yue) Simplified Chinese (zh-CN) Dataset Structure JSON lines with yue field and zh field… See the full description on the dataset page: https://huggingface.co/datasets/botisan-ai/cantonese-mandarin-translations.texttranslation10K<n<100K31 likes124 downloads3y agoHugging Face20MagicDataTech /magicdata-dialect-cantonese-tts-lite MagicData-Dialect-Northeastern Chinese-TTS-Lite MAGIC DATA OPEN-SOURCE LICENSE Dataset Overview Item Information Dataset Type N/A Language Chinese Dialect Speech Style Scripted Content N/A Audio Parameters 48 kHz, 16 bits File Format WAV (PCM) Recording Equipment microphone Recording Environment quiet indoor environment License Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicDataTech/magicdata-dialect-cantonese-tts-lite.audion<1K0 likes124 downloads4mo agoHugging Face21alvanlii /cantonese-youtube-transcriptiontext10K<n<100K8 likes108 downloads2y agoHugging Face22mesolitica /Cantonese-Radio-Description-Instructions Cantonese-Radio-Description-Instructions Originally from alvanlii/cantonese-radio, we use Qwen/Qwen2.5-72B-Instruct to generate description based on the transcription. how to prepare the dataset huggingface-cli download \ mesolitica/Cantonese-Radio-Description-Instructions \ --include '*.zip' \ --repo-type "dataset" \ --local-dir './' wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Cantonese-Radio-Description-Instructions.audio100K<n<1M0 likes106 downloads1y agoHugging Face23raptorkwok /cantonese_sentencesThis dataset contains raw Cantonese sentences obtained from Hong Kong-based local forum. Some sentences are processed and used in the Cantonese-Written Chinese Parallel Corpus Gen 3. text10M<n<100M7 likes105 downloads10mo agoHugging Face24HKAllen /cantonese-chinese-parallel-corpus Dataset Summary This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation. The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation. Languages Cantonese (yue) Simplified Chinese (zh) Dataset Structure Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page: https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.texttranslation100K<n<1M3 likes101 downloads2y agoHugging Face25wcyat /shikoto-cantonese-sweettextn<1K1 likes98 downloads2y agoHugging Face26Pangeanic /Cantonese-Japanese-Machine-Translation-Corpus-text PangeanicYueJa - Cantonese Japanese Parallel Corpus PangeanicYueJa is a Cantonese-Japanese parallel corpus designed for machine translation, multilingual large language model (LLM) training, cross-lingual NLP research, retrieval-augmented generation (RAG), bilingual embeddings, instruction tuning, and multilingual AI systems. This release contains 55,000 Cantonese-Japanese sentence pairs sampled from a larger corpus of approximately 3.08 million parallel sentence pairs. For the… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Cantonese-Japanese-Machine-Translation-Corpus-text.texttranslation10K<n<100K4 likes97 downloads3mo agoHugging Face27open-llm-leaderboard-old /details_hon9kon9ize__CantoneseLLM-6B-preview202402 Dataset Card for Evaluation run of hon9kon9ize/CantoneseLLM-6B-preview202402 Dataset automatically created during the evaluation run of model hon9kon9ize/CantoneseLLM-6B-preview202402 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_hon9kon9ize__CantoneseLLM-6B-preview202402.0 likes80 downloads3y agoHugging Face28Pangeanic /Cantonese-English-Machine-Translation-Corpus-text PangeanicYueEn - Cantonese English Parallel Corpus PangeanicYueEn is a large-scale Cantonese-English parallel corpus designed for machine translation, multilingual large language model (LLM) training, cross-lingual NLP research, retrieval-augmented generation (RAG), bilingual embeddings, instruction tuning, and multilingual AI systems. This release contains 150,000 Cantonese-English sentence pairs sampled from a larger corpus of approximately 8.05 million parallel sentence… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Cantonese-English-Machine-Translation-Corpus-text.texttranslation100K<n<1M1 likes80 downloads3mo agoHugging Face29ziyou-li /cantonese_processed_guangzhou Dataset Card for "cantonese_processed_guangzhou" More Information needed 1K<n<10K0 likes78 downloads4y agoHugging Face30JackyHoCL /cleaned_mixed_cantonese_and_english_speechAll right reserved by and credit to AlienKevin/mixed_cantonese_and_english_speech This is a cleaned verison from AlienKevin/mixed_cantonese_and_english_speech: https://huggingface.co/datasets/AlienKevin/mixed_cantonese_and_english_speech Removed '"' in the preffix and suffix Removed empty records in order to reduce hallucination ------2025-08-03------ Converted to MP3, reduce size to 1/10 of the original size. audio10K<n<100K4 likes77 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.