CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01raptorkwok /cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences. texttranslation100K<n<1M17 likes194 downloads3y agoHugging Face02botisan-ai /cantonese-mandarin-translations Dataset Card for cantonese-mandarin-translations Dataset Summary This is a machine-translated parallel corpus between Cantonese (a Chinese dialect that is mainly spoken by Guangdong (province of China), Hong Kong, Macau and part of Malaysia) and Chinese (written form, in Simplified Chinese). Supported Tasks and Leaderboards N/A Languages Cantonese (yue) Simplified Chinese (zh-CN) Dataset Structure JSON lines with yue field and zh field… See the full description on the dataset page: https://huggingface.co/datasets/botisan-ai/cantonese-mandarin-translations.texttranslation10K<n<100K31 likes124 downloads3y agoHugging Face03HKAllen /cantonese-chinese-parallel-corpus Dataset Summary This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation. The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation. Languages Cantonese (yue) Simplified Chinese (zh) Dataset Structure Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page: https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.texttranslation100K<n<1M3 likes101 downloads2y agoHugging Face04him0413 /cantonese-qa-instructions 🇭🇰 Cantonese QA Instructions (v0.3) 粵語 / 廣東話指令微調數據集 — 全合成、全 QC'd、全繁體中文輸出 A high-quality synthetic instruction-tuning dataset of natural spoken Cantonese queries paired with Traditional Chinese answers (50–200 characters). Covers 6 diverse domains at varying difficulty levels. Generated by Qwen 3.6 Dense and quality-controlled by DeepSeek V4 Pro. Fully automated nightly generation pipeline on dedicated hardware. 🔗 View on Hugging Face 📊 Dataset Stats (v0.3)… See the full description on the dataset page: https://huggingface.co/datasets/him0413/cantonese-qa-instructions.textquestion-answering1K<n<10K0 likes72 downloads3mo agoHugging Face05raptorkwok /cantonese-written-chinese-translationtexttranslation1K<n<10K1 likes53 downloads2y agoHugging Face06stvlynn /Cantonese-Dialoguetext10K<n<100K11 likes36 downloads2y agoHugging Face07zeno109 /cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences. texttranslation100K<n<1M1 likes35 downloads4mo agoHugging Face08ReopenAI /cantonese-youtube-transcription-fusionhttps://huggingface.co/datasets/alvanlii/cantonese-youtube 数据集中train-00000-of-01090.parquet 到 train-00350-of-01090.parquet 部分的转写文本清洗。使用qwen3-asr、qwen3-omni、sensevoicesmall(https://huggingface.co/ASLP-lab/WSYue-ASR) 进行转写,然后用 Qwen3.6-35B-A3B 根据语义进行转写纠正。 text100K<n<1M0 likes33 downloads1mo agoHugging Face09agentlans /cantonese-chinese Cantonese-Mandarin-Traditional Chinese Parallel Corpus This dataset provides a parallel corpus of Cantonese, Simplified Chinese, and Traditional Chinese text. Dataset Composition The dataset is a combination of two existing datasets: botisan-ai/cantonese-mandarin-translations raptorkwok/cantonese-chinese-dataset-gen2 Train Set: Merged from both source datasets Test and Validation Sets: Derived from raptorkwok/cantonese-chinese-dataset-gen2 Language Variants… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/cantonese-chinese.texttranslation1M<n<10M2 likes23 downloads2y agoHugging Face10cantonesesra /Cantonese_AllAspectQA_11K Cantonese_AllAspectQA_11K A comprehensive Question-Answer dataset in Cantonese (粵語) covering a wide range of conversational topics and aspects. Overview Cantonese_AllAspectQA_11K is a curated collection of 11,000 question-answer pairs in Cantonese, designed to facilitate the development, training, and evaluation of Cantonese language models and conversational AI systems. The dataset captures authentic Cantonese speech patterns, colloquialisms, and cultural nuances across… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_AllAspectQA_11K.texttext-generation10K<n<100K3 likes21 downloads1y agoHugging Face11cantonesesra /Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K Yue_WizardLMEvolved_AllAspectQA_Small_1.5K A specialized collection of high-quality question-answer pairs in Cantonese (粵語) inspired by the WizardLM evolution methodology, covering diverse and complex topics. Overview Yue_WizardLMEvolved_AllAspectQA_Small_1.5K is a curated dataset of 1,500 evolved question-answer pairs in Cantonese. This dataset applies the WizardLM evolution philosophy to generate in-depth, nuanced responses to complex questions in Cantonese. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K.texttext-generation1K<n<10K0 likes19 downloads1y agoHugging Face12Nin8520 /Cantonese_QAQA datasets implemented with simple diversification extensions to word datasets textquestion-answering100K<n<1M3 likes15 downloads1y agoHugging Face13open-llm-leaderboard /hon9kon9ize__CantoneseLLMChat-v0.5-detailsgated Dataset Card for Evaluation run of hon9kon9ize/CantoneseLLMChat-v0.5 Dataset automatically created during the evaluation run of model hon9kon9ize/CantoneseLLMChat-v0.5 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/hon9kon9ize__CantoneseLLMChat-v0.5-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face14open-llm-leaderboard /lordjia__Llama-3-Cantonese-8B-Instruct-detailsgated Dataset Card for Evaluation run of lordjia/Llama-3-Cantonese-8B-Instruct Dataset automatically created during the evaluation run of model lordjia/Llama-3-Cantonese-8B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lordjia__Llama-3-Cantonese-8B-Instruct-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face15wangrongsheng /Cantonese-Datatext10K<n<100K0 likes9 downloads2y agoHugging Face16open-llm-leaderboard /hon9kon9ize__CantoneseLLMChat-v1.0-7B-detailsgated Dataset Card for Evaluation run of hon9kon9ize/CantoneseLLMChat-v1.0-7B Dataset automatically created during the evaluation run of model hon9kon9ize/CantoneseLLMChat-v1.0-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/hon9kon9ize__CantoneseLLMChat-v1.0-7B-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face17pendingremove32894 /cantonese_allaspectqa_11ktext10K<n<100K0 likes8 downloads2y agoHugging Face18open-llm-leaderboard /lordjia__Qwen2-Cantonese-7B-Instruct-detailsgated Dataset Card for Evaluation run of lordjia/Qwen2-Cantonese-7B-Instruct Dataset automatically created during the evaluation run of model lordjia/Qwen2-Cantonese-7B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lordjia__Qwen2-Cantonese-7B-Instruct-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face19andyloyola /cantonese-gba-foundrygated Cantonese + GBA Synthetic Data High-quality synthetic dialogues for Hong Kong / Greater Bay Area business scenarios. Commercial License: HK$20000 one-time (full commercial rights) How to buy: Click "Request Access" Write "I want commercial license" I will send you a Stripe invoice immediately Pay → instant full download Generated with a proprietary hybrid AI pipeline for maximum cultural accuracy and natural Cantonese-English code-switching. textn<1K0 likes2 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.