datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.AICC🔧 🔧 Our New-Gen Html Parser MinerU-HTML Now Realease!
AICC: AI-ready Common Crawl Dataset
Paper | Project page
News
[2025-12-24] 🔥 CC-MinerU-Code Updated! We have updated our specialized high-quality code dataset CC-MinerU-Code, containing 4.58M samples, also extracted from the full Common Crawl corpus.
Download: CC-MinerU-Code
Each record includes language, code_language, and Markdown-formatted content with fenced code blocks. Here is a sample:
{… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/AICC.SlimPajama-Meta-rater
Annotated SlimPajama Dataset
Dataset Description
This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions.
Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.awesome-markdown-ebooks
Awesome-markdown-ebooks
Your GitHub PDFs, Now AI-Ready.
Project repo: https://github.com/OpenDataLab/awesome-markdown-ebooks
Spark-234K
Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature
🎉 Accepted to EMNLP 2026 Findings!
Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/Spark-234K.Open-MOPD-Data
Open-MOPD Data
This repository contains the training and evaluation data released with
Open-MOPD, including mixed-domain supervised fine-tuning data, the shared
RL/OPD prompt mixture, and six evaluation benchmarks.
Dataset contents
Configuration
Description
Examples
rl_prompt_mix
Shared math, code, and instruction-following prompts for RL and OPD
86,931
sft_openr1_math_93k
Math SFT data in a unified think-tag format
93,733
sft_ocr_50k
Sampled… See the full description on the dataset page: https://huggingface.co/datasets/BytedTsinghua-SIA/Open-MOPD-Data.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.DILA_OPENDATA_FR_2023
French Government Open Data (DILA) Dataset - 2023
Overview
The French Government Open Data (DILA) Dataset is a collection of text data extracted from various sources provided by the French government, specifically the Direction de l'information légale et administrative (DILA). This dataset contains a wide range of legal, administrative, and legislative documents. The data has been organized into several categories for easy access and analysis.
Dataset Splits… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/DILA_OPENDATA_FR_2023.MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking
MMFineReason-SFT-123K
The Hardest 7% — Less Data, More Reasoning
📖 Overview
MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0).
🎯 Key Highlights
123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.ocs-krisdika
Open Law Data Thailand: OCS Krisdika Dataset
ชุดข้อมูลกฎหมายจาก สำนักงานคณะกรรมการกฤษฎีกา (Office of the Council of State) รวบรวมและจัดทำโดยโครงการ Open Law Data Thailand เพื่อส่งเสริมการเข้าถึงข้อมูลกฎหมายในรูปแบบที่เครื่องอ่านได้ (Machine-Readable)
Dataset Structure
ข้อมูลถูกจัดเก็บในรูปแบบ JSON Lines (.jsonl) แบ่งไฟล์ตาม ปีและเดือน (YYYY/YYYY-MM.jsonl) เพื่อความสะดวกในการดาวน์โหลดและบริหารจัดการ
Data Fields
แต่ละบรรทัด (Row) ประกอบด้วยข้อมูลดังนี้:
title… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/ocs-krisdika.ODA-Math-460k
ODA-Math-460k
ODA-Math-460k is a large-scale math reasoning dataset curated from top-performing open mathematics corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination, LLM-based filtering, and verifier-backed response distillation.It targets a “learnable but challenging” difficulty band: non-trivial for smaller models yet solvable by stronger reasoning models.
🧠 Dataset Summary
Domain: Mathematics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Math-460k.SlimPajama-Meta-rater-Readability-30B
Top 30B token SlimPajama Subset selected by the Readability rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.open_r1_dataset
Integration of publicly available datasets related to R1
We integrate all data and remove contaminated data and data with inconsistent formats. The user defaults to selecting version 'V1', with a total of 2592286 samples.
1 Relevant datasets mentioned in HuggingFace/open_r1:
(1) HuggingFaceH4/numina-deepseek-r1-qwen-7b: A dataset distilled using DeepSeek-R1-Distill-Qwen-7B.
Hugging Face downloads: 631.
(2) AI-MO/NuminaMath-TIR: A subset of 70K math-related samples… See the full description on the dataset page: https://huggingface.co/datasets/xiushenghuang/open_r1_dataset.MathLake
MathLake: A Large-Scale Mathematics Dataset
MathLake is a massive collection of 8.3 million mathematical problems aggregated from over 50 open-source datasets. Unlike datasets focused on filtering for the highest quality solutions immediately, MathLake prioritizes query comprehensiveness, serving as a universal "raw ore" for researchers to curate, distill, or annotate further. The dataset provides annotations for Difficulty, Format, and Subject, covering diverse mathematical fields… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MathLake.ODA-Fin-SFT-318k
Unlocking Data Value in Finance: A Study on Distillation
and Difficulty-Aware Training
📖 Overview
ODA-Fin-SFT-318K is a meticulously curated financial reasoning dataset comprising 318,599 samples with high-quality Chain-of-Thought (CoT) annotations. Constructed via multi-stage distillation from Qwen3-235B-A22B-Thinking and rigorous verification, this dataset establishes a robust foundation for training financial language models with strong reasoning capabilities.… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Fin-SFT-318k.SlimPajama-Meta-rater-Professionalism-30B
Top 30B token SlimPajama Subset selected by the Professionalism rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Professionalism dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness)… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Professionalism-30B.open-web-math-minhash
Dataset Card for "open-web-math-minhash"
An attempt at a "high quality sample" of open-web-math/open-web-math by aggressively applying minhash from text-dedup. The result is 1.82M rows down from the original 6M:
DatasetDict({
train: Dataset({
features: ['url', 'text', 'date', 'metadata'],
num_rows: 1820241
})
})
Usage
Unless you need the metadata, load the text-only config which is only 1.4 GB/5 shards:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/open-web-math-minhash.SlimPajama-Meta-rater-Reasoning-30B
Top 30B token SlimPajama Subset selected by the Reasoning rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Reasoning dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework. Each… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Reasoning-30B.QUEST-SFT-Data-Open-ended
QUEST SFT Data (Open-ended)
Project Page | Paper | GitHub
Open-ended supervised fine-tuning trajectories for QUEST (tool-using assistant format). Split: train. Columns: messages (list[{role, content}]).
Load
from datasets import load_dataset
ds = load_dataset("osunlp/QUEST-SFT-Data-Open-ended", split="train", streaming=True)
row = next(iter(ds))
print(row.keys())
QUEST Family
Type
Resources
35B checkpoints
RL, MT+SFT, MT, SFT
30B checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/QUEST-SFT-Data-Open-ended.WanJuan-Korean
💡 Introduction
WanJuan-Korean(万卷丝路-韩语) corpus, with a volume exceeding 280GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Korean.ODA-Fin-RL-12k
Unlocking Data Value in Finance: A Study on Distillation
and Difficulty-Aware Training
📖 Overview
ODA-Fin-RL-12K is a carefully curated dataset for reinforcement learning (RL) in financial domain, comprising 12,187 hard-but-verifiable samples. Designed to complement ODA-Fin-SFT-318K, this dataset targets challenging financial reasoning tasks with concise, reliably verifiable answers—optimized for RL training.
🎯 Key Highlights
12K Hard Samples: Curated… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Fin-RL-12k.open-christian-data
Open Christian Data
Open Christian Data (OCD) aims to be the single unified collection of all public domain Christian text. It exists to bring this collection together from across the internet and to structure it in a useful format for public use.
Beyond the Bible, Christian writing is poorly represented as a cohesive dataset or as data structured for AI training. This Hugging Face release is the AI-focused publication of the collection: consistent, downloadable JSON for model… See the full description on the dataset page: https://huggingface.co/datasets/OpenChristianDataOrg/open-christian-data.Open_Reaction_Data
ORDerly: Styrene Mizoroki-Heck RAG-Ready Dataset
This repository contains chemical reaction data formatted for Retrieval-Augmented Generation (RAG) systems.
The data is a processed version of the ORDerly benchmark, specifically focusing on reaction conditions and forward/retro prediction tasks.
Dataset Structure
The data is split into 10,000-row Parquet chunks to prevent Out-of-Memory (OOM) errors during ingestion into vector databases.
It includes:
orderly_condition:… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Open_Reaction_Data.SlimPajama-Meta-rater-Cleanliness-30B
Top 30B token SlimPajama Subset selected by the Cleanliness rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Cleanliness dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Cleanliness-30B.WanJuan-Thai
💡 Introduction
WanJuan-Thai (万卷丝路-泰语) corpus, with a volume exceeding 155GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Thai.open-math-dataset
Dataset Description
Open Math Dataset is an open mathematics corpus designed for mathematical AI, reasoning, education, and research.
The project is being developed from Sri Lanka with the goal of creating an internationally useful mathematics dataset for developers, researchers, educators, and AI systems.
Mathematics Corpus
The dataset is designed to contain structured mathematical problems and solutions across different mathematical domains and education levels.… See the full description on the dataset page: https://huggingface.co/datasets/tharustack/open-math-dataset.SA-Prot-annot
SA-Prot-Annot Dataset (Sci-Align)
🌌 The Sciverse Data Foundation
Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research.
Sciverse consists of three core data… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SA-Prot-annot.WanJuan-BaiHua
OpenDataLab近期拟发布万卷·百华大规模专业领域数据集,主要面向金融、能源、文化教育、政务、通信、交通运输、医疗健康、汽车、烟草、计算机等领域大模型训练,提供高质量、精细处理、领域分类的多模态专用语料。 在此我们将面向社区征集需求,若有相关行业数据集的需求,请填写以下调查问卷,我们会根据社区反馈来决定不同领域数据集的发布顺序,欢迎大家提供想法!
调查问卷链接:https://www.wjx.cn/vm/mxip3eE.aspx#
数据集样例详见:https://opendatalab.com/OpenDataLab/WanJuan-BaiHua
K12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究
