datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chinese-ai-detection-dataset
Chinese AI Detection Dataset
中文AI文本检测数据集
数据集简介
用于训练中文AI生成文本检测模型的综合数据集,包含纯人类、纯AI以及混合文本(人类+AI)。
核心特色:使用[SEP]标记显式标注混合文本的人类/AI边界。
数据统计
类型
样本数
说明
总计
66,001
训练/验证/测试集
纯人类
27,719
多领域人类文本
纯AI
27,719
多模型生成
C2 (续写)
3,781
人类开头+AI续写
C3 (改写)
3,781
AI改写人类文本
C4 (润色)
3,001
AI润色人类文本
数据格式
{
"text": "文本内容(混合文本包含[SEP]标记)",
"label": 0, // 0=Human, 1=AI
"category": "C2", // Human/AI/C2/C3/C4
"source": "数据来源"
}… See the full description on the dataset page: https://huggingface.co/datasets/AnxForever/chinese-ai-detection-dataset.chinese-stock-datasetChinese-Student-English-Essay
Dataset Card for Chinese Student English Essay (CSEE) Dataset
Dataset Summary
The Chinese Student English Essay (CSEE) dataset is designed for Automated Essay Scoring (AES) tasks. It consists of 13,270 English essays written by high school students in Beijing, who are English as a Second Language (ESL) learners. These essays were collected from two final exams and correspond to two writing prompts.
Each essay is evaluated across three key dimensions by experienced… See the full description on the dataset page: https://huggingface.co/datasets/Xiaochr/Chinese-Student-English-Essay.chinese-official-source-reachability
Reachability of Chinese official verification portals from outside China
If you are doing due diligence on a Chinese supplier, the advice is always "check the official registry". This dataset measures whether you can.
Eight official Chinese verification sources were measured from public vantage points outside mainland China across several rounds between August and September 2026, plus two controls (www.gov.cn and www.baidu.com).
The result is not "Chinese government sites are… See the full description on the dataset page: https://huggingface.co/datasets/derrick459/chinese-official-source-reachability.chinese_lyricsnumbers-in-chinese-poetry
Numerals in Classical Chinese Poetry (Tang–Qing)
Zizheng Lv · ORCID 0009-0004-0327-4477 · DOI 10.5281/zenodo.22734274
Code and documentation: https://github.com/zizhenglvubc/numbers-in-chinese-poetry
Dataset summary
Numeral counts for 175,355 classical Chinese poems from the Tang dynasty through
the Qing. Two files:
poem_level_numerals.csv.gz — one row per poem, 175,355 rows.
numeral_rates_by_group.csv — one row per collection, 7 rows including a total.
59.24%… See the full description on the dataset page: https://huggingface.co/datasets/zizhenglvubc/numbers-in-chinese-poetry.Multi-Doc-Multi-QA-ChineseDeprecated, please use Multi-Doc-QA-Chinese instead.
文档和问答对都来自 Multi-Doc-QA-Chinese,通过随机抽取和组合形成多轮问答形式。
推荐直接使用原始数据集Multi-Doc-QA-Chinese自己生成指令微调数据,可以控制参考文档和问答的数量
经过随机组合,每条数据形成了 20-60个参考文档 + 10个问答对的形式
chat格式为chatml
chinese-jokebenchmark-1-chinese-m2mInfo:
Translated on Chinese by facebook/m2m100_418M model
Source: jayavibhav/prompt-injection-safety
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Chinese translated by facebook/m2m100_418M
score_zh_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-chinese-m2m.benchmark-4-chinese-m2mInfo:
Translated on Chinese by facebook/m2m100_418M model
Source: nvidia/Aegis-AI-Content-Safety-Dataset-2.0
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 1,000 prompts (500… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-4-chinese-m2m.scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.chinese_poetryChinese-Student-English-Essay
Dataset Card for Chinese Student English Essay (CSEE) Dataset
Dataset Summary
The Chinese Student English Essay (CSEE) dataset is designed for Automated Essay Scoring (AES) tasks. It consists of 13,270 English essays written by high school students in Beijing, who are English as a Second Language (ESL) learners. These essays were collected from two final exams and correspond to two writing prompts.
Each essay is evaluated across three key dimensions by experienced… See the full description on the dataset page: https://huggingface.co/datasets/zhongshiting/Chinese-Student-English-Essay.benchmark-4-chinese-gtInfo:
Translated on Chinese by Google Translate
Source: nvidia/Aegis-AI-Content-Safety-Dataset-2.0
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 1,000 prompts (500 safe / 500… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-4-chinese-gt.benchmark-1-chinese-gtInfo:
Translated on Chinese by Google Translate
Source: jayavibhav/prompt-injection-safety
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Chinese translated by Google Translate
score_zh_google - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-chinese-gt.benchmark-2-chinese-gtInfo:
Translated on Chinese by Google Translate
Source: xTRam1/safe-guard-prompt-injection
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Chinese translated by Google Translate
score_zh_google - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-2-chinese-gt.benchmark-2-chinese-m2mInfo:
Translated on Chinese by facebook/m2m100_418M model
Source: xTRam1/safe-guard-prompt-injection
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Chinese translated by facebook/m2m100_418M
score_zh_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-2-chinese-m2m.benchmark-3-chinese-m2mInfo:
Translated on Chinese by facebook/m2m100_418M model
Source: JailbreakBench/JBB-Behaviors
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 200 prompts (100 safe / 100… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-3-chinese-m2m.ChineseToEngbenchmark-3-chinese-gtInfo:
Translated on Chinese by Google Translate
Source: JailbreakBench/JBB-Behaviors
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 200 prompts (100 safe / 100 unsafe)… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-3-chinese-gt.chinese-outbound-relocation-index-2026
Chinese Outbound Relocation Index 2026
Open-data ranking of 30 destinations for mainland Chinese and Hong Kong emigrants, weighted for education, safety, healthcare, wealth mobility, visa pathways, and 2025-2026 policy flags.
Files
chinese-outbound-relocation-index-2026.csv: CSV export for Chinese Outbound Relocation Index 2026.
chinese-outbound-relocation-index-2026.json: JSON API response for Chinese Outbound Relocation Index 2026.
Canonical Source… See the full description on the dataset page: https://huggingface.co/datasets/Walker7777/chinese-outbound-relocation-index-2026.chinese_chaste_tree_lagundi_dataset
