datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences.
cantonese-mandarin-translations
Dataset Card for cantonese-mandarin-translations
Dataset Summary
This is a machine-translated parallel corpus between Cantonese (a Chinese dialect that is mainly spoken by Guangdong (province of China), Hong Kong, Macau and part of Malaysia) and Chinese (written form, in Simplified Chinese).
Supported Tasks and Leaderboards
N/A
Languages
Cantonese (yue)
Simplified Chinese (zh-CN)
Dataset Structure
JSON lines with yue field and zh field… See the full description on the dataset page: https://huggingface.co/datasets/botisan-ai/cantonese-mandarin-translations.cantonese-chinese-parallel-corpus
Dataset Summary
This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation.
The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation.
Languages
Cantonese (yue)
Simplified Chinese (zh)
Dataset Structure
Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page: https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.cantonese-qa-instructions
🇭🇰 Cantonese QA Instructions (v0.3)
粵語 / 廣東話指令微調數據集 — 全合成、全 QC'd、全繁體中文輸出
A high-quality synthetic instruction-tuning dataset of natural spoken Cantonese queries paired with Traditional Chinese answers (50–200 characters). Covers 6 diverse domains at varying difficulty levels. Generated by Qwen 3.6 Dense and quality-controlled by DeepSeek V4 Pro. Fully automated nightly generation pipeline on dedicated hardware.
🔗 View on Hugging Face
📊 Dataset Stats (v0.3)… See the full description on the dataset page: https://huggingface.co/datasets/him0413/cantonese-qa-instructions.cantonese-written-chinese-translationCantonese-Dialoguecantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences.
cantonese-youtube-transcription-fusionhttps://huggingface.co/datasets/alvanlii/cantonese-youtube 数据集中train-00000-of-01090.parquet 到 train-00350-of-01090.parquet 部分的转写文本清洗。使用qwen3-asr、qwen3-omni、sensevoicesmall(https://huggingface.co/ASLP-lab/WSYue-ASR)
进行转写,然后用 Qwen3.6-35B-A3B 根据语义进行转写纠正。
cantonese-chinese
Cantonese-Mandarin-Traditional Chinese Parallel Corpus
This dataset provides a parallel corpus of Cantonese, Simplified Chinese, and Traditional Chinese text.
Dataset Composition
The dataset is a combination of two existing datasets:
botisan-ai/cantonese-mandarin-translations
raptorkwok/cantonese-chinese-dataset-gen2
Train Set: Merged from both source datasets
Test and Validation Sets: Derived from raptorkwok/cantonese-chinese-dataset-gen2
Language Variants… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/cantonese-chinese.Cantonese_AllAspectQA_11K
Cantonese_AllAspectQA_11K
A comprehensive Question-Answer dataset in Cantonese (粵語) covering a wide range of conversational topics and aspects.
Overview
Cantonese_AllAspectQA_11K is a curated collection of 11,000 question-answer pairs in Cantonese, designed to facilitate the development, training, and evaluation of Cantonese language models and conversational AI systems. The dataset captures authentic Cantonese speech patterns, colloquialisms, and cultural nuances across… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_AllAspectQA_11K.Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K
Yue_WizardLMEvolved_AllAspectQA_Small_1.5K
A specialized collection of high-quality question-answer pairs in Cantonese (粵語) inspired by the WizardLM evolution methodology, covering diverse and complex topics.
Overview
Yue_WizardLMEvolved_AllAspectQA_Small_1.5K is a curated dataset of 1,500 evolved question-answer pairs in Cantonese. This dataset applies the WizardLM evolution philosophy to generate in-depth, nuanced responses to complex questions in Cantonese. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K.Cantonese_QAQA datasets implemented with simple diversification extensions to word datasets
hon9kon9ize__CantoneseLLMChat-v0.5-details
Dataset Card for Evaluation run of hon9kon9ize/CantoneseLLMChat-v0.5
Dataset automatically created during the evaluation run of model hon9kon9ize/CantoneseLLMChat-v0.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/hon9kon9ize__CantoneseLLMChat-v0.5-details.lordjia__Llama-3-Cantonese-8B-Instruct-details
Dataset Card for Evaluation run of lordjia/Llama-3-Cantonese-8B-Instruct
Dataset automatically created during the evaluation run of model lordjia/Llama-3-Cantonese-8B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lordjia__Llama-3-Cantonese-8B-Instruct-details.Cantonese-Datahon9kon9ize__CantoneseLLMChat-v1.0-7B-details
Dataset Card for Evaluation run of hon9kon9ize/CantoneseLLMChat-v1.0-7B
Dataset automatically created during the evaluation run of model hon9kon9ize/CantoneseLLMChat-v1.0-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/hon9kon9ize__CantoneseLLMChat-v1.0-7B-details.cantonese_allaspectqa_11klordjia__Qwen2-Cantonese-7B-Instruct-details
Dataset Card for Evaluation run of lordjia/Qwen2-Cantonese-7B-Instruct
Dataset automatically created during the evaluation run of model lordjia/Qwen2-Cantonese-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lordjia__Qwen2-Cantonese-7B-Instruct-details.cantonese-gba-foundry
Cantonese + GBA Synthetic Data
High-quality synthetic dialogues for Hong Kong / Greater Bay Area business scenarios.
Commercial License: HK$20000 one-time (full commercial rights)
How to buy:
Click "Request Access"
Write "I want commercial license"
I will send you a Stripe invoice immediately
Pay → instant full download
Generated with a proprietary hybrid AI pipeline for maximum cultural accuracy and natural Cantonese-English code-switching.
