datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chinese-tts-text-10m
Chinese Conversational TTS Text
Synthetic everyday spoken Mandarin (zh-CN) utterances for TTS synthesis and training.
Each row is a single, standalone conversational turn — the kind of thing a person actually
says at home, at work, or with friends — paired with an emotion label and a ready-to-use
style prompt.
Built to drive two engines:
Qwen3-TTS — via the instruct column, which feeds its natural-language style channel.
Zonos2 (expressive) — via spoken_emotion, for matching… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/chinese-tts-text-10m.chinese_text_correction
Dataset Card
中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。
Repository: shibing624/pycorrector
Dataset Summary
拼写纠错数据
lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2
ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data
medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition
chinese-iching-divination-text
Note
This dataset contains information extracted from the following ancient Chinese books. Please note that this dataset is limited by specific limitations. For more robust models, consider using additional datasets or data augmentation techniques.
数据集包含以下古代中国书籍中的信息。
易经-49部
连山易-清-马国翰.txt
周易口义-宋-胡瑗.txt
周易举正-唐-郭京.txt
易原-清-多隆阿.txt
新本郑氏周易-清-恵栋.txt
伊川易传-宋-程颐.txt
易纂言外翼洛书说-元-吴澄.txt
周易正义-唐-孔颖达.txt
易纬略义-清-张惠言.txt
易筮通变-元-雷思齐.txt
易经证释-清-陆宗舆.txt
易纬乾元序制记-汉-郑玄.txt
易纬辨终备-汉-郑玄.txt… See the full description on the dataset page: https://huggingface.co/datasets/pokkoa/chinese-iching-divination-text.chinese-street-text-baiduTaken from here: https://aistudio.baidu.com/datasetdetail/8429 ; This solely acts as a nice way to integrate it to my scripts
balanced_SROIE_CHINESE_IAM_text_recognitionchinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition
chinese-poetry-text
Dataset Card for "chinese-poetry-text"
More Information needed
chinese-iching-divination-text
Note
This dataset contains information extracted from the following ancient Chinese books. Please note that this dataset is limited by specific limitations. For more robust models, consider using additional datasets or data augmentation techniques.
数据集包含以下古代中国书籍中的信息。
易经-49部
连山易-清-马国翰.txt
周易口义-宋-胡瑗.txt
周易举正-唐-郭京.txt
易原-清-多隆阿.txt
新本郑氏周易-清-恵栋.txt
伊川易传-宋-程颐.txt
易纂言外翼洛书说-元-吴澄.txt
周易正义-唐-孔颖达.txt
易纬略义-清-张惠言.txt
易筮通变-元-雷思齐.txt
易经证释-清-陆宗舆.txt
易纬乾元序制记-汉-郑玄.txt
易纬辨终备-汉-郑玄.txt… See the full description on the dataset page: https://huggingface.co/datasets/Evanwayne/chinese-iching-divination-text.task775_pawsx_chinese_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task775_pawsx_chinese_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task775_pawsx_chinese_text_modification.LLM-Chinese-Textual-Disambiguation
Chinese Textual Ambiguity Dataset
This dataset is the accompanying dataset for the paper:
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
Paper (arXiv): https://arxiv.org/abs/2507.23121
Project repository: https://github.com/ictup/LLM-Chinese-Textual-Disambiguation
Dataset Summary
This release contains 925 Chinese textual ambiguity records collected and annotated for research on ambiguity detection, ambiguity understanding… See the full description on the dataset page: https://huggingface.co/datasets/pip1237/LLM-Chinese-Textual-Disambiguation.Chinese_Text_False_DataNYUAD-ComNets/Chinese_Text_False_Data was used as adversarial attacks structured as few lines Chinese Text about the headline.
Example:
拜登-哈里斯政府发布的国防部指令5420.01引发了对公民自由和对美国公民使用武力的重大担忧。
该指令允许在抗议者面前潜在使用致命武力,这在应对社会动荡时可能被视为极端措施,
特别是在2024年总统选举前政治紧张局势加剧的背景下。批评者认为,这种政策削弱了言论和集会的基本权利,
可能加剧执法部门与平民之间的冲突。支持者可能认为这是在动荡情况下维持秩序和保护公共安全的必要步骤。
该指令的影响可能对政府与公民之间的关系以及美国抗议权的更广泛讨论产生持久影响。
This dataset was used in the paper titled "Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial… See the full description on the dataset page: https://huggingface.co/datasets/NYUAD-ComNets/Chinese_Text_False_Data.Chinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat
Chinese-English Parallel Translation Corpus (Chinese Source Text & English Translation)
A Chinese–English parallel corpus resource for translation and cross-lingual alignment applications, providing one-to-one bilingual text pairs: Chinese source texts aligned with their corresponding English translations. The data covers common writing styles and domains, making it suitable for parallel alignment, translation modeling, and cross-lingual representation learning.
It supports… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Chinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat.chinese_text_correction
Dataset Card
中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。
Repository: shibing624/pycorrector
Dataset Summary
拼写纠错数据
lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2
ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data
medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/xlp100/chinese_text_correction.flan_combined_task775_pawsx_chinese_text_modificationChinese_Textvietnamese_chinese_ancient_text-40kclassical-chinese-text-generation
