datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese-LiPS
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
⭐ Introduction
The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios.
🚀 Dataset Details
Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.Traditional-Chinese-Medicine-Multiple_choice_question
Discription
This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.new-title-chineseChinese-Psychology-Books
免责声明与使用须知 (Disclaimer and Usage Notice)
数据集内容
本数据集包含从互联网上多个来源收集的 中文心理学电子书 的集合。
许可证
本数据集的组织结构、汇编方式以及由维护者添加的任何元数据或注释根据 知识共享署名-非商业性使用 4.0 国际许可协议 (Creative Commons Attribution-NonCommercial 4.0 International License - CC BY-NC 4.0) 提供。这意味着您可以基于非商业目的分享和修改这部分内容,但必须给出适当的署名。
请注意:此 CC BY-NC 4.0 许可证不适用于数据集中包含的原始电子书文件本身。
版权声明
数据集中包含的个别电子书文件极有可能受到版权法保护,其版权归各自的作者、出版商或其他版权所有者所有。
数据集维护者不拥有这些电子书的版权。
这些电子书的来源多样且零散,部分来源可能难以追溯。
使用限制与责任… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Psychology-Books.Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.chinese_text_correction
Dataset Card
中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。
Repository: shibing624/pycorrector
Dataset Summary
拼写纠错数据
lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2
ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data
medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.chinese-american-elder-fraud-qa
chinese-american-elder-fraud-qa
A hand-authored, trilingual (Mandarin / Cantonese / English) fraud-recognition dataset for first-generation Chinese-American elders and the adult children who help them. 235 rows authored, 207 adapted through the Adaption Labs platform with reasoning traces. Grounded in FBI, IC3, and SFPD reports on Chinese-community elder fraud.
Adaption Labs Uncharted Data Challenge submission.
Metric
Value
Rows authored
235
Rows adapted (training… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/chinese-american-elder-fraud-qa.Traditional-Chinese-Medicine-Exam
Coming Soon...
chinese-ai-detection-dataset
Chinese AI Detection Dataset
中文AI文本检测数据集
数据集简介
用于训练中文AI生成文本检测模型的综合数据集,包含纯人类、纯AI以及混合文本(人类+AI)。
核心特色:使用[SEP]标记显式标注混合文本的人类/AI边界。
数据统计
类型
样本数
说明
总计
66,001
训练/验证/测试集
纯人类
27,719
多领域人类文本
纯AI
27,719
多模型生成
C2 (续写)
3,781
人类开头+AI续写
C3 (改写)
3,781
AI改写人类文本
C4 (润色)
3,001
AI润色人类文本
数据格式
{
"text": "文本内容(混合文本包含[SEP]标记)",
"label": 0, // 0=Human, 1=AI
"category": "C2", // Human/AI/C2/C3/C4
"source": "数据来源"
}… See the full description on the dataset page: https://huggingface.co/datasets/AnxForever/chinese-ai-detection-dataset.chinese-stock-datasetChineseOCRBenchChinese-Student-English-Essay
Dataset Card for Chinese Student English Essay (CSEE) Dataset
Dataset Summary
The Chinese Student English Essay (CSEE) dataset is designed for Automated Essay Scoring (AES) tasks. It consists of 13,270 English essays written by high school students in Beijing, who are English as a Second Language (ESL) learners. These essays were collected from two final exams and correspond to two writing prompts.
Each essay is evaluated across three key dimensions by experienced… See the full description on the dataset page: https://huggingface.co/datasets/Xiaochr/Chinese-Student-English-Essay.chinese-official-source-reachability
Reachability of Chinese official verification portals from outside China
If you are doing due diligence on a Chinese supplier, the advice is always "check the official registry". This dataset measures whether you can.
Eight official Chinese verification sources were measured from public vantage points outside mainland China across several rounds between August and September 2026, plus two controls (www.gov.cn and www.baidu.com).
The result is not "Chinese government sites are… See the full description on the dataset page: https://huggingface.co/datasets/derrick459/chinese-official-source-reachability.ChineseEnglishTranslationDatasetChinese_sentiment45k_python_code_chinese_instruction
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
中文提示的代码数据集
其中提示部分通过调用GPT-4.0-turbo API翻译成中文
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/jean1/45k_python_code_chinese_instruction.chinese-logic-sentiment-dataset
中文逻辑情感分析数据集 (由豆包API生成)
-->English
这是一个专门用于增强中文情感分析模型逻辑推理能力的数据集,由 豆包 API 生成并经过人工清洗/筛选。
它包含 反讽(Irony)、双重否定(Double Negative)、转折(Transition) 和 简单句(Simple) 四种逻辑类型。
数据结构
该数据集包含三个部分:
train.csv: 训练集,包含 2176 条样本。
val.csv: 验证集,包含 545 条样本。
test.csv: 测试集,包含 960 条样本。用于模型最终评估。
数据字段
Field
Description
text
中文文本内容
label
情感标签,0 表示负面,1 表示正
type
逻辑类型,包含反讽、双重否定、转折和简单句
sub_type
逻辑子类型,进一步细分逻辑结构
domain
领域 (影视/美食/旅游/生活/购物/社交)
使用指南
from… See the full description on the dataset page: https://huggingface.co/datasets/YiMeng-SYSU/chinese-logic-sentiment-dataset.Chinese-Speech-Dataset
🎧 Chinese (Simplified) Speech Dataset
The Chinese (Simplified) speech dataset is a high-quality speech audio dataset developed to support scalable AI and machine learning solutions with diverse and structured audio data. It contains 105 hours of speech data across 700 audio files, delivered in MP3 and WAV formats, with a total size of 229 MB. This well-balanced audio dataset provides reliable voice data, featuring 54% female and 46% male speakers, with age distribution ranging from… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Chinese-Speech-Dataset.GLM-Open-Dialogue-Chinese-Dataset-v1Simplified_Chinese_Multi-Emotion_Dialogue_Dataset
Simplified_Chinese_Multi-Emotion_Dialogue_Dataset
数据说明
本数据是简体中文口语情感分类数据集
翻译自于:Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset
使用Qwen2.5-32B-instruce模型将其从繁体中文翻译为简体中文
共4159条
情感类别
原始类别
翻译后类别
条数
悲傷語調
伤心
486
憤怒語調
生气
527
關切語調
关心
560
驚奇語調
惊讶
499
開心語調
开心
592
平淡語氣
平静
705
厭惡語調
厌恶
404
chinese_conversation_and_spam
Caution! This dataset contains explicit language and fraud information. Use at your own risk!
For AutoTrain use: please select Text Classification (Binary) as Task.
What is included
conversations in chinese under tag 0
spam conversations under tag1
Where does the data come from
part of the data came from conversations in Chinese Telegram groups
part of them are from logging channels of anti-spam bots
How many data is included
A total of 9.9k… See the full description on the dataset page: https://huggingface.co/datasets/paulkm/chinese_conversation_and_spam.h2o-translated-chinese-med-prompts
Translated Chinese Medical Prompts
This repository contains medical prompts translated originally from Chinese, which can be used as training data for natural language processing (NLP) tasks related to the medical domain in English language.
Dataset Description
The dataset consists of a collection of medical prompts originally in Chinese, which have been translated into English. These prompts cover various medical topics, including symptoms, diagnoses, treatments, medications, and… See the full description on the dataset page: https://huggingface.co/datasets/h2oai/h2o-translated-chinese-med-prompts.GLM-Open-Dialogue-Chinese-Dataset-v2chinese_word_frequency
Chinese Word frequency statistics
word segmentation by jieba tool
seq_monkey_data : statistics on 13,000,000 document / 6,561,241,266 words / 11,313,242,610 characters
sft-chinesechinese-sensitive-topics-qa
Chinese Sensitive Topics QA Dataset
Dataset Summary
This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.chinese_lyricsnumbers-in-chinese-poetry
Numerals in Classical Chinese Poetry (Tang–Qing)
Zizheng Lv · ORCID 0009-0004-0327-4477 · DOI 10.5281/zenodo.22734274
Code and documentation: https://github.com/zizhenglvubc/numbers-in-chinese-poetry
Dataset summary
Numeral counts for 175,355 classical Chinese poems from the Tang dynasty through
the Qing. Two files:
poem_level_numerals.csv.gz — one row per poem, 175,355 rows.
numeral_rates_by_group.csv — one row per collection, 7 rows including a total.
59.24%… See the full description on the dataset page: https://huggingface.co/datasets/zizhenglvubc/numbers-in-chinese-poetry.GLM-Open-Dialogue-Chinese-DatasetChinese_Recipe_small
