datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SVAMPchinese-fineweb-edu
This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 !
Chinese Fineweb Edu Dataset [中文] [English]
[OpenCSG Community] [👾github] [wechat] [Twitter]
📖Technical Report
Chinese Fineweb Edu dataset is a meticulously constructed high-quality Chinese pre-training corpus, specifically designed for natural language processing tasks in the education domain. This dataset undergoes a rigorous selection and… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu.fineweb-edu
📚 FineWeb-Edu
1.3 trillion tokens of the finest educational data the 🌐 web has to offer
Paper: https://arxiv.org/abs/2406.17557
What is it?
📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version.
To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/fineweb-edu.ChineseWebText2.0-HighQuality
📘 ChineseWebText2.0-HighQuality
Overview
ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original
CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License).
This subset retains only samples with:
quality_score ≥ 0.9
toxicity.score ≤ 0.01
The goal is to provide a cleaner and more reliable dataset suitable for
language model pre-training, instruction tuning, and quality-sensitive downstream tasks.
This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.chinese-fineweb-edu-v2
This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 !
Chinese Fineweb Edu Dataset V2 [中文] [English]
[OpenCSG Community] [👾github] [wechat] [Twitter]
📖Technical Report
Chinese Fineweb Edu Dataset V2 is a comprehensive upgrade of the original Chinese Fineweb Edu, designed and optimized for natural language processing (NLP) tasks in the education sector. This high-quality Chinese pretraining dataset has… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu-v2.smoltalk-chinese
Chinese SmolTalk Dataset [中文] [English]
[OpenCSG Community] [👾github] [wechat] [Twitter]
📖Technical Report
smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/smoltalk-chinese.Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier.
Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make.
(͡ ° ͜ʖ ͡ °)
tags: socks, garter belt, foot fetish, ntr, netori.....
Thanks Moleys/Numeron for the dataset donation.
universal_spanish_chilean_corpus
Universal Chilean Spanish Corpus
Este dataset se compone de 37_213_992 textos correspondientes a español de Chile y a español multidialectal.
Los textos en español multidialectal provienen del spanish books.
Los textos en español de Chile vienen de los dominios .cl del mc4 dataset y de tweets, noticias y reclamos de l chilean-spanish-corpus
Name
Count
Source
books
87967
spanish books
mc4
8706681
from mc4 (.cl domains) in chilean-spanish-corpus
twitter
27306583… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/universal_spanish_chilean_corpus.Chinese-Instruct-Lite
中文指令微调数据集 - Lite 版本
💻 Github Repo
[!TIP]
这不是 Chinese-Instruct 的子集,而是一个全新的简化数据集。
如果您想要一个可以真实使用、而不仅仅适用于学习的数据集,欢迎访问:Mxode/Chinese-Instruct
如果您想要一个更加简单易收敛、主题集中的数据集,可以访问:Mxode/I_Wonder_Why-Chinese
具体构成
本数据集包含如下 5 个子集,总数据量 10M+。
code:代码主题的指令数据集,数据量 1.2M+。
math:数学主题的指令数据集,数据量 1.7M+。
general:通用指令数据集,主题广泛,与 code 和 math 指令不重复,数据量 5.1M+。
math(reasoning):数学推理数据集,指令采样自 math 子集,可通过 id 关联,数据量 1.2M+。code(reasoning):代码推理数据集,指令采样自 code 子集,可通过 id 关联,数据量 700K+。
如何使用… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Instruct-Lite.c4-chinese-zhtw
Dataset Card for "c4-chinese-zhtw"
內容
Common Crawl 是一個非營利組織,負責抓取網路並向公眾免費提供其檔案和資料集。Common Crawl 的網路檔案包含自 2008 年以來收集的 PB 級資料。它一般每月完成一次抓取。
Common Crawl 的爬蟲程式遵守 nofollow 和 robots.txt 政策。用於處理 Common Crawl 資料集的開源程式碼是公開可用的。
這個繁中的數據來是來自 Common Crawl 2023-14 的 data archive 下載并進行清理 。
這是 jed351 準備的版本,託管在這個位址:
https://huggingface.co/datasets/jed351/Traditional-Chinese-Common-Crawl-Filtered
支援的任務
C4主要用於預訓練語言模型(pretrain language model)。
範例
一個樣本的範例:
{… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/c4-chinese-zhtw.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.OpenOrca-Traditional-Chinese🐋 OpenOrca-Chinese 数据集!🐋
感謝 Open-Orca/OpenOrca 資料集的發布,為廣大NLP研究人員和開發者帶來了寶貴的資源!
這是一個對 Open-Orca/OpenOrca 資料集中文翻譯的版本,翻譯引擎為 Google 翻譯,希望能為中文 LLM 研究做出一點點貢獻。
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/lchakkei/OpenOrca-Traditional-Chinese.I_Wonder_Why-Chinese
🧐 十万个为什么 - 中文百科开放问答数据集
💻 Github Repo
这是一个中文百科开放问答数据集,共分为 3 个子集:general、preference、reasoning。这个数据集可适用于 SFT 指令微调、DPO 类强化学习、R1 类推理蒸馏任务。
[!tip]
[2025/05/09] 发布了一个新的中文指令数据集 Chinese-Instruct-Lite,包含代码、数学、通用多场景,同样包含一般指令微调数据与推理数据,数据总量 10M+
[2025/05/05] 更新:数据集扩增,现在指令由 600K+ 增加到 1.2M+ 了!
数据集详情
所有的子集共享相同的指令(prompt),共计 1.2M+,每一条指令都有自己独有的 12 位 id。这意味着你可以根据 id 交叉混合使用不同的子集。
由于指令相同,因此所有子集的数据量都是一致的,均为 1.2M+。
general:这个子集适用于 SFT 指令微调,形式是最简单的 prompt-response 格式。… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/I_Wonder_Why-Chinese.Chinese_Debate_Documents
Dataset Card for Chinese Debate Documents
ASR-transcribed corpus of competitive Mandarin Chinese university debates,
speaker-segmented and timestamped, with topic / round / team metadata parsed
from the source filenames.
Loading the Dataset
from datasets import load_dataset
ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train")
print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"])
for seg in ds[0]["segments"][:3]:
print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.china-effective-laws-regulations
全国现行法律法规合集
现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。
数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。
这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。
快照日期:2026-08-26
效力说明
本数据集 以现行有效法律法规为主体:
效力 status
法规份数
说明
有效
17,649
现行有效,默认应使用这一部分
尚未生效
7
已公布、施行日晚于快照日
失效
45
文件名含「失效」,多为已到期的全国人大常委会试点授权决定
使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.clean_passages_80m-chinese-zhtw
Dataset Card for "clean_passages_80m-chinese-zhtw"
包含8千萬餘萬(88328203)個中文段落,不包含任何字母、數字。文字長度大部分介於 50~200 個字。
原始資料集是用於訓練GENIUS模型中文版。論文參考引用:
@article{guo2022genius,
title={GENIUS: Sketch-based Language Model Pre-training via Extreme and Selective Masking for Text Generation and Augmentation},
author={Guo, Biyang and Gong, Yeyun and Shen, Yelong and Han, Songqiao and Huang, Hailiang and Duan, Nan and Chen, Weizhu},
journal={arXiv preprint arXiv:2211.10330},
year={2022}
}… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/clean_passages_80m-chinese-zhtw.OpenOrca-Chinese🐋 OpenOrca-Chinese 数据集!🐋
感谢 Open-Orca/OpenOrca 数据集的发布,给广大NLP研究人员和开发者带来了宝贵的资源!
这是一个对 Open-Orca/OpenOrca 数据集中文翻译的版本,翻译引擎为 Google 翻译,希望能给中文 LLM 研究做出一点点贡献。
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/yys/OpenOrca-Chinese.TC260-Chinese-Safety-Prompts
TC260 Chinese Safety Prompts V1
Public research dataset containing synthetic Chinese safety-testing prompts.
Records have different quality tiers; the full dataset must not be described
as human-verified or Gold data.
这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据
由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径
扫描、精确去重和四字shingle近似去重。
本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。
类别名称和映射用于研究性实现,不构成法律、监管或合规结论。
数据规模
原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。
A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBBBBBBQ/TC260-Chinese-Safety-Prompts.securecode-web-archive
SecureCode Web: Traditional Web & Application Security Dataset
Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance
Paper | GitHub | Dataset | Model Collection | Blog Post
What's new in v2.6
v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had
shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-web-archive.en-vocab-mnemonicsThis dataset is produced as a part of my Capstone (see Github). The dataset uses roughly 500 examples from Balepur et. al, 2024.
Shallow-Encoding Mnemonics
Deep-Encoding Mnemonics
Homophonic: olfactory sounds like "old factory."
Etymology: preposterous - pre (before) + post (after) + erous, which implies absurd.
Chunking: obsequious sounds like "ob-se-ki-ass. Obedient servants kiss your ass
Morphology: Suffixes "ate" are usually verbs. Prefix "ab" means from, away.
Keyword:… See the full description on the dataset page: https://huggingface.co/datasets/chiffonng/en-vocab-mnemonics.CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact but high-difficulty synthetic reasoning datasetwith long Chain-of-Thought (CoT) trajectories and broad STEM coverage, designed for reasoning post-training. All examples are fully LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
🔥 Why CHIMERA?
Recent reasoning advances rely heavily on high-quality… See the full description on the dataset page: https://huggingface.co/datasets/TianHongZXY/CHIMERA.en-vocab-en-mnemonics-cotchinese-writing-benchmark
Zhiyin: Exploring the Frontier of Chinese LLM Writing
Website • GitHub • Hugging Face
Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks.
Benchmark Overview
Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5.
Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-benchmark.Fineweb-Edu-Chinese-V2.1
Chinese Fineweb Edu Dataset V2.1 [中文] [English]
[OpenCSG Community] [👾github] [wechat] [Twitter]
📖Technical Report
The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different folders… See the full description on the dataset page: https://huggingface.co/datasets/alphagocc/Fineweb-Edu-Chinese-V2.1.SWE-smith-trajectories
SWE-smith Trajectories
Code
•
Paper
•
Site
This dataset contains the 5017 trajectories we fine-tuned Qwen 2.5 Coder Instruct on, leading to
SWE-agent-LM-32B, a coding LM agent that
achieve 40.2% on SWE-bench Verified (no verifiers or multiple rollouts, just 1 attempt per instance).
Trajectories were generated by running SWE-agent + Claude 3.7 Sonnet on task instances from
the SWE-smith dataset.
openorca-chinese-zhtw
Dataset Card for "openorca-chinese-zhtw"
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope.
The data is primarily used for training and evaluation in the field of natural… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/openorca-chinese-zhtw.ielts-writing-task2-essays
📚 IELTS Writing Task 2 Essays & Feedback Dataset (Writing9)
Dataset Summary
This dataset contains 8,000+ real IELTS Writing Task 2 essays crawled from Writing9. It covers 128 real IELTS exam questions categorized into 25 topics (such as Art, Business, Education, Technology, Environment, Government, Health, etc.).
Each record includes:
essay_id: Unique identifier on Writing9
topic: Topic category (e.g. Art, Business and Companies, Cities)
question: Cleaned IELTS… See the full description on the dataset page: https://huggingface.co/datasets/chillies/ielts-writing-task2-essays.smoltalk-chinese
Chinese SmolTalk Dataset [中文] [English]
[OpenCSG Community] [👾github] [wechat] [Twitter]
📖Technical Report
smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/sunorme/smoltalk-chinese.reasoning-sft-CHIMERA
reasoning-sft-CHIMERA
Converted version of TianHongZXY/CHIMERA, filtered and reformatted for SFT/reasoning training. Both subsets (Qwen3-235B-2507 and Qwen3.5-397B) are included. No content was modified or regenerated, just reformatted the columns into a standard messages format.
Filtering
Kept only rows with correctness == True
Randomly dropped 50% of Mathematics rows to reduce math dominance
Both subsets combined into a single file
Format
Each row has three… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-CHIMERA.chinese-writing-bench-judgements-gpt-5.4
Zhiyin: Exploring the Frontier of Chinese LLM Writing
Website • GitHub • Hugging Face
Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks.
Benchmark Overview
Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5.
Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements-gpt-5.4.
