datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Fineweb-Edu-Chinese-V2.2
Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train)
[[中文]] | [[English]]
OpenCSG Community | 👾 GitHub | 📖 Technical Report
Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs
Chinese Fineweb Edu Dataset V2.2is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain.
This project aims to solve the core pain point of… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.2.ShareGPT-Chinese-English-90k
ShareGPT-Chinese-English-90k Bilingual Human-Machine QA Dataset
A high-quality Chinese-English parallel bilingual human-machine QA dataset, covering user questions in real and complex scenarios. It is used for training high-quality dialogue models (more robust in instruction distribution than those datasets generated by repeatedly calling API interfaces to simulate machine-generated Q&A, like Moss)
Features:
Provides fully semantically equivalent Chinese-English parallel corpus… See the full description on the dataset page: https://huggingface.co/datasets/shareAI/ShareGPT-Chinese-English-90k.doctrina-jurisprudencia-chile
🇨🇱 Corpus Jurídico y Doctrinal de Chile en Markdown (Open Legal Chile)
Bienvenido al repositorio oficial del Corpus Jurídico Canónico, Doctrinal y Jurisprudencial de Chile, desarrollado y mantenido por Open Legal Chile.
Este repositorio ofrece acceso 100% completo, libre y gratuito (Apache-2.0) al texto íntegro de la dogmática jurídica chilena, a las Guías Oficiales de la Academia Judicial, a los fallos de los Tribunales Ambientales (1TA, 2TA, 3TA), sus anuarios y boletines, y… See the full description on the dataset page: https://huggingface.co/datasets/pablobenavidesj/doctrina-jurisprudencia-chile.Chinese-SimpleQA
Overview
🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper • 📊 Leaderboard
Chinese SimpleQA is the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, our benchmark covers 6 major topics with 99 diverse subtopics.
Please visit our website or check our paper for more details.… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-SimpleQA.Traditional_Chinese-aya_collection
資料集描述
繁體中文 Aya (Traditional Chinese Aya Chinese;TCA):專注於繁體中文處理的 Aya 集合的精選子集
概述
繁體中文 Aya 是一個精心策劃的資料集,源自 CohereForAI 的綜合 Aya 集合,特別關注繁體中文文本資料。
此資料集結合了來自 CohereForAI/aya_collection,過濾掉除繁體中文、簡體中文內容之外的所有內容。
目標
繁體中文 Aya 的目標是為研究人員、技術專家和語言學家提供即用型繁體中文文本資源,顯著減少專注於繁體中文的 NLP 和 AI 專案中數據預處理所需的時間和精力。
資料集來源與資訊
資料來源: 從 CohereForAI/aya_collection 64 個子集而來。
語言: 繁體中文、簡體中文('zho')
應用: 非常適合語言建模、文本分類、情感分析、和機器翻譯等任務。
論文連結: 2402.06619
維護人: Heng666
License: Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/Heng666/Traditional_Chinese-aya_collection.Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier.
Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make.
(͡ ° ͜ʖ ͡ °)
tags: socks, garter belt, foot fetish, ntr, netori.....
Thanks Moleys/Numeron for the dataset donation.
Magpie-Qwen2-Pro-200K-Chinese
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2-Pro-200K-Chinese.HC3-ChineseHuman ChatGPT Comparison Corpus (HC3) Chinese VersionFineweb-Edu-Chinese-V2.2
Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train)
[[中文]] | [[English]]
OpenCSG Community | 👾 GitHub | 📖 Technical Report
Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs
Chinese Fineweb Edu Dataset V2.2is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain.
This project aims to solve the core pain point of… See the full description on the dataset page: https://huggingface.co/datasets/enche1561/Fineweb-Edu-Chinese-V2.2.Chinese-Instruct-Lite
中文指令微调数据集 - Lite 版本
💻 Github Repo
[!TIP]
这不是 Chinese-Instruct 的子集,而是一个全新的简化数据集。
如果您想要一个可以真实使用、而不仅仅适用于学习的数据集,欢迎访问:Mxode/Chinese-Instruct
如果您想要一个更加简单易收敛、主题集中的数据集,可以访问:Mxode/I_Wonder_Why-Chinese
具体构成
本数据集包含如下 5 个子集,总数据量 10M+。
code:代码主题的指令数据集,数据量 1.2M+。
math:数学主题的指令数据集,数据量 1.7M+。
general:通用指令数据集,主题广泛,与 code 和 math 指令不重复,数据量 5.1M+。
math(reasoning):数学推理数据集,指令采样自 math 子集,可通过 id 关联,数据量 1.2M+。code(reasoning):代码推理数据集,指令采样自 code 子集,可通过 id 关联,数据量 700K+。
如何使用… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Instruct-Lite.Chinese-DeepSeek-R1-Distill-data-110k
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。
该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.Traditional-Chinese-Medicine-Multiple_choice_question
Discription
This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集
Wizard-LM包含了很多难度超过Alpaca的指令。
中文的问题翻译会有少量指令注入导致翻译失败的情况
中文回答是根据中文问题再进行问询得到的。
我们会陆续将更多数据集发布到hf,包括
Coco Caption的中文翻译
CoQA的中文翻译
CNewSum的Embedding数据
增广的开放QA数据
WizardLM的中文翻译
如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。
骆驼(Luotuo): 开源中文大语言模型
https://github.com/LC1332/Luotuo-Chinese-LLM
骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。
( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 )
骆驼项目不是商汤科技的官方产品。
Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.Chinese-Instruct
中文指令微调数据集
💻 Github Repo
本项目旨在构建一个高质量、多领域、大规模的中文指令微调数据集。
本项目将会持续更新。更多数据集欢迎访问 Github Repo。
[!TIP]
如果您想要一个可用于学习的简化版中文指令数据集,可以访问:Mxode/Chinese-Instruct-Lite
具体构成
dpsk-r1-distil:中文 DeepSeek-R1 蒸馏数据集,来自 Congliu/Chinese-DeepSeek-R1-Distill-data-110k,根据打分质量做了筛选,提取了最终的回答,未包含思考过程。
chinese-reasoning-distil:中文推理蒸馏数据集,来自 Mxode/Chinese-Reasoning-Distil-Data,提取了最终的回答,未包含思考过程。
firefly:中文通用指令微调数据集,指令取自 Mxode/Firefly-1.1M-Rephrased,其本身已经相较于原 Firefly… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Instruct.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.OpenOrca-Traditional-Chinese🐋 OpenOrca-Chinese 数据集!🐋
感謝 Open-Orca/OpenOrca 資料集的發布,為廣大NLP研究人員和開發者帶來了寶貴的資源!
這是一個對 Open-Orca/OpenOrca 資料集中文翻譯的版本,翻譯引擎為 Google 翻譯,希望能為中文 LLM 研究做出一點點貢獻。
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/lchakkei/OpenOrca-Traditional-Chinese.I_Wonder_Why-Chinese
🧐 十万个为什么 - 中文百科开放问答数据集
💻 Github Repo
这是一个中文百科开放问答数据集,共分为 3 个子集:general、preference、reasoning。这个数据集可适用于 SFT 指令微调、DPO 类强化学习、R1 类推理蒸馏任务。
[!tip]
[2025/05/09] 发布了一个新的中文指令数据集 Chinese-Instruct-Lite,包含代码、数学、通用多场景,同样包含一般指令微调数据与推理数据,数据总量 10M+
[2025/05/05] 更新:数据集扩增,现在指令由 600K+ 增加到 1.2M+ 了!
数据集详情
所有的子集共享相同的指令(prompt),共计 1.2M+,每一条指令都有自己独有的 12 位 id。这意味着你可以根据 id 交叉混合使用不同的子集。
由于指令相同,因此所有子集的数据量都是一致的,均为 1.2M+。
general:这个子集适用于 SFT 指令微调,形式是最简单的 prompt-response 格式。… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/I_Wonder_Why-Chinese.china-effective-laws-regulations
全国现行法律法规合集
现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。
数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。
这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。
快照日期:2026-08-26
效力说明
本数据集 以现行有效法律法规为主体:
效力 status
法规份数
说明
有效
17,649
现行有效,默认应使用这一部分
尚未生效
7
已公布、施行日晚于快照日
失效
45
文件名含「失效」,多为已到期的全国人大常委会试点授权决定
使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.OpenOrca-Chinese🐋 OpenOrca-Chinese 数据集!🐋
感谢 Open-Orca/OpenOrca 数据集的发布,给广大NLP研究人员和开发者带来了宝贵的资源!
这是一个对 Open-Orca/OpenOrca 数据集中文翻译的版本,翻译引擎为 Google 翻译,希望能给中文 LLM 研究做出一点点贡献。
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/yys/OpenOrca-Chinese.chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.Chinese-DeepSeek-R1-Distill-data-110k-SFT
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:该版本为,可以直接SFT使用的版本,将原始数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。该中文数据集中的数据分布如下:
Math:共计36568个样本,
Exam:共计2432个样本,
STEM:共计12648个样本,… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT.perceptual-constancy
Perceptual Constancy
Perceptual Constancy is a multimodal benchmark designed to evaluate high-level perceptual invariance in large vision-language models (VLMs). It probes a model’s understanding of physical and geometric stability under varying sensory appearances. This dataset is part of the Grow AI Like a Child benchmark initiative.
🧠 Dataset Overview
The Perceptual Constancy dataset focuses on appearance-invariant reasoning using both static images and short… See the full description on the dataset page: https://huggingface.co/datasets/grow-ai-like-a-child/perceptual-constancy.Fineweb-Edu-Chinese-V3
Fineweb-Edu-Chinese-V3
中文 | English
OpenCSG 社区 | 数据集许可协议
数据集简介
Fineweb-Edu-Chinese-V3 是 OpenCSG 面向学科知识问答、教材理解和推理型指令微调场景构建的高质量中英双语教育 SFT 数据集,也是 Fineweb-Edu-Chinese 系列的最新版本。
该版本包含 18.81 万条 SFT 样本,来自 100,442 篇高质量图书、教材、学科文献与技术长文,覆盖计算机、自然科学、社科人文、法学、经济五大学科方向,并同步提供 Messages、Messages-no-system、Alpaca 三种训练格式。三种格式是同一批问答对的不同导出视图,训练时应按模型模板选择其中一种,而不是简单相加作为独立数据规模。
V3 是 Fineweb-Edu-Chinese 系列的一次数据源与构造范式的整体切换。V1.0 至 V2.3 均以大规模中文网页语料为基础:通过打分器筛选出具备教育属性的网页文本,再由大模型生成问答。V3… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V3.Fineweb-Edu-Chinese-V2.3
Chinese Fineweb Edu Dataset V2.3
中文 | English
OpenCSG 社区 | GitHub | 数据集许可协议
数据集简介
Chinese Fineweb Edu Dataset V2.3 是 OpenCSG 面向中文教育、知识问答、指令微调和文本生成场景构建的高质量中文教育 SFT 数据集。
该版本包含 23.04 万条高质量中文教育 QA pairs,并将同一批问答对发布为 Alpaca、Messages、Messages-no-system 三种训练格式。三种格式面向不同训练模板,建议训练时按模型和框架选择其中一种格式使用,而不是将不同格式简单相加作为独立知识规模。
V2.3 是在 V2.2 基础上的质量升级版本。针对 V2.2 社区反馈和内部质量审计中出现的重复模式、异常中英文混入、噪声片段、弱证据支撑回答和低质量合成输出等问题,V2.3 提高了源文本进入生成环节的门槛,并优化了问答生成与过滤逻辑。
在数据构建上,V2.3 从约 2.3T… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.3.chinese-clean-energy-battery-open-intelligence
🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset
Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.chinese-ai-and-robotics-open-intelligence
🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset
Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.securecode-web-archive
SecureCode Web: Traditional Web & Application Security Dataset
Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance
Paper | GitHub | Dataset | Model Collection | Blog Post
What's new in v2.6
v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had
shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-web-archive.CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact but high-difficulty synthetic reasoning datasetwith long Chain-of-Thought (CoT) trajectories and broad STEM coverage, designed for reasoning post-training. All examples are fully LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
🔥 Why CHIMERA?
Recent reasoning advances rely heavily on high-quality… See the full description on the dataset page: https://huggingface.co/datasets/TianHongZXY/CHIMERA.Chinese-High-School-Chemistry-Correction-Dataset
Chinese-High-School-Chemistry-Correction-Dataset
一个面向「高中化学垂直大模型微调」的中文问答与文本生成数据集
1. 数据集缘起
为了训练一个高中化学领域的垂直大模型,我们需要大量高质量、结构化的中文语料。本数据集整理了三版主流教科书、常考化学方程式与畅销教辅等中的知识点,全部转为统一的 JSONL 格式。
2. 数据来源
普通高中教科书(苏教版、人教版、鲁教版)、高中常考化学方程式、高中参考教辅资料(一本涂书、教材帮等)均转成jsonl格式
该jsonl文件数据,部分行或许有格式错误,需要自行编写py脚本校对,以便用于大模型微调。
3. 数据格式(JSONL)
每行一条记录,可直接用于 Hugging Face datasets 库:
{"instruction": "已知0.5 mol的水(H₂O)的质量是9 g,且含有3.01×10²³个水分子。请计算1 mol水的质量和阿伏伽德罗常数。", "output":… See the full description on the dataset page: https://huggingface.co/datasets/liushuaiqian/Chinese-High-School-Chemistry-Correction-Dataset.
