Chinese text
chinese-tts-text-10m
Chinese Conversational TTS Text
Synthetic everyday spoken Mandarin (zh-CN) utterances for TTS synthesis and training.
Each row is a single, standalone conversational turn — the kind of thing a person actually
says at home, at work, or with friends — paired with an emotion label and a ready-to-use
style prompt.
Built to drive two engines:
Qwen3-TTS — via the instruct column, which feeds its natural-language style channel.
Zonos2 (expressive) — via spoken_emotion, for matching… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/chinese-tts-text-10m.chinese_text_correction
Dataset Card
中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。
Repository: shibing624/pycorrector
Dataset Summary
拼写纠错数据
lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2
ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data
medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition
chinese-iching-divination-text
Note
This dataset contains information extracted from the following ancient Chinese books. Please note that this dataset is limited by specific limitations. For more robust models, consider using additional datasets or data augmentation techniques.
数据集包含以下古代中国书籍中的信息。
易经-49部
连山易-清-马国翰.txt
周易口义-宋-胡瑗.txt
周易举正-唐-郭京.txt
易原-清-多隆阿.txt
新本郑氏周易-清-恵栋.txt
伊川易传-宋-程颐.txt
易纂言外翼洛书说-元-吴澄.txt
周易正义-唐-孔颖达.txt
易纬略义-清-张惠言.txt
易筮通变-元-雷思齐.txt
易经证释-清-陆宗舆.txt
易纬乾元序制记-汉-郑玄.txt
易纬辨终备-汉-郑玄.txt… See the full description on the dataset page: https://huggingface.co/datasets/pokkoa/chinese-iching-divination-text.Chinese-Image-Text-Corpus-dataset
REILX/Chinese-Image-Text-Corpus-dataset
[ English | 中文 ]
Introduction
The REILX/Chinese-Image-Text-Corpus-dataset is a multimodal dataset that pairs Chinese textual data with corresponding images. This dataset is derived from the Chinese-Xinhua Dictionary Database, which includes idioms, single characters, words, and aphorisms.
Dataset Structure
The dataset is organized into the following categories:
Idioms: Traditional Chinese idioms with explanations and… See the full description on the dataset page: https://huggingface.co/datasets/REILX/Chinese-Image-Text-Corpus-dataset.chinese-street-text-baiduTaken from here: https://aistudio.baidu.com/datasetdetail/8429 ; This solely acts as a nice way to integrate it to my scripts
