datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cantonese-qa-instructions
🇭🇰 Cantonese QA Instructions (v0.3)
粵語 / 廣東話指令微調數據集 — 全合成、全 QC'd、全繁體中文輸出
A high-quality synthetic instruction-tuning dataset of natural spoken Cantonese queries paired with Traditional Chinese answers (50–200 characters). Covers 6 diverse domains at varying difficulty levels. Generated by Qwen 3.6 Dense and quality-controlled by DeepSeek V4 Pro. Fully automated nightly generation pipeline on dedicated hardware.
🔗 View on Hugging Face
📊 Dataset Stats (v0.3)… See the full description on the dataset page: https://huggingface.co/datasets/him0413/cantonese-qa-instructions.Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K
Yue_WizardLMEvolved_AllAspectQA_Small_1.5K
A specialized collection of high-quality question-answer pairs in Cantonese (粵語) inspired by the WizardLM evolution methodology, covering diverse and complex topics.
Overview
Yue_WizardLMEvolved_AllAspectQA_Small_1.5K is a curated dataset of 1,500 evolved question-answer pairs in Cantonese. This dataset applies the WizardLM evolution philosophy to generate in-depth, nuanced responses to complex questions in Cantonese. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K.Cantonese_AllAspectQA_11K
Cantonese_AllAspectQA_11K
A comprehensive Question-Answer dataset in Cantonese (粵語) covering a wide range of conversational topics and aspects.
Overview
Cantonese_AllAspectQA_11K is a curated collection of 11,000 question-answer pairs in Cantonese, designed to facilitate the development, training, and evaluation of Cantonese language models and conversational AI systems. The dataset captures authentic Cantonese speech patterns, colloquialisms, and cultural nuances across… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_AllAspectQA_11K.
