datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AICC🔧 🔧 Our New-Gen Html Parser MinerU-HTML Now Realease!
AICC: AI-ready Common Crawl Dataset
Paper | Project page
News
[2025-12-24] 🔥 CC-MinerU-Code Updated! We have updated our specialized high-quality code dataset CC-MinerU-Code, containing 4.58M samples, also extracted from the full Common Crawl corpus.
Download: CC-MinerU-Code
Each record includes language, code_language, and Markdown-formatted content with fenced code blocks. Here is a sample:
{… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/AICC.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.SlimPajama-Meta-rater
Annotated SlimPajama Dataset
Dataset Description
This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions.
Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.Spark-234K
Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature
🎉 Accepted to EMNLP 2026 Findings!
Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/Spark-234K.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.Open-MOPD-Data
Open-MOPD Data
This repository contains the training and evaluation data released with
Open-MOPD, including mixed-domain supervised fine-tuning data, the shared
RL/OPD prompt mixture, and six evaluation benchmarks.
Dataset contents
Configuration
Description
Examples
rl_prompt_mix
Shared math, code, and instruction-following prompts for RL and OPD
86,931
sft_openr1_math_93k
Math SFT data in a unified think-tag format
93,733
sft_ocr_50k
Sampled… See the full description on the dataset page: https://huggingface.co/datasets/BytedTsinghua-SIA/Open-MOPD-Data.DILA_OPENDATA_FR_2023
French Government Open Data (DILA) Dataset - 2023
Overview
The French Government Open Data (DILA) Dataset is a collection of text data extracted from various sources provided by the French government, specifically the Direction de l'information légale et administrative (DILA). This dataset contains a wide range of legal, administrative, and legislative documents. The data has been organized into several categories for easy access and analysis.
Dataset Splits… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/DILA_OPENDATA_FR_2023.MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking
MMFineReason-SFT-123K
The Hardest 7% — Less Data, More Reasoning
📖 Overview
MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0).
🎯 Key Highlights
123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.ODA-Math-460k
ODA-Math-460k
ODA-Math-460k is a large-scale math reasoning dataset curated from top-performing open mathematics corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination, LLM-based filtering, and verifier-backed response distillation.It targets a “learnable but challenging” difficulty band: non-trivial for smaller models yet solvable by stronger reasoning models.
🧠 Dataset Summary
Domain: Mathematics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Math-460k.open_r1_dataset
Integration of publicly available datasets related to R1
We integrate all data and remove contaminated data and data with inconsistent formats. The user defaults to selecting version 'V1', with a total of 2592286 samples.
1 Relevant datasets mentioned in HuggingFace/open_r1:
(1) HuggingFaceH4/numina-deepseek-r1-qwen-7b: A dataset distilled using DeepSeek-R1-Distill-Qwen-7B.
Hugging Face downloads: 631.
(2) AI-MO/NuminaMath-TIR: A subset of 70K math-related samples… See the full description on the dataset page: https://huggingface.co/datasets/xiushenghuang/open_r1_dataset.SlimPajama-Meta-rater-Readability-30B
Top 30B token SlimPajama Subset selected by the Readability rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.MathLake
MathLake: A Large-Scale Mathematics Dataset
MathLake is a massive collection of 8.3 million mathematical problems aggregated from over 50 open-source datasets. Unlike datasets focused on filtering for the highest quality solutions immediately, MathLake prioritizes query comprehensiveness, serving as a universal "raw ore" for researchers to curate, distill, or annotate further. The dataset provides annotations for Difficulty, Format, and Subject, covering diverse mathematical fields… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MathLake.open-web-math-minhash
Dataset Card for "open-web-math-minhash"
An attempt at a "high quality sample" of open-web-math/open-web-math by aggressively applying minhash from text-dedup. The result is 1.82M rows down from the original 6M:
DatasetDict({
train: Dataset({
features: ['url', 'text', 'date', 'metadata'],
num_rows: 1820241
})
})
Usage
Unless you need the metadata, load the text-only config which is only 1.4 GB/5 shards:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/open-web-math-minhash.SlimPajama-Meta-rater-Professionalism-30B
Top 30B token SlimPajama Subset selected by the Professionalism rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Professionalism dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness)… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Professionalism-30B.SlimPajama-Meta-rater-Reasoning-30B
Top 30B token SlimPajama Subset selected by the Reasoning rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Reasoning dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework. Each… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Reasoning-30B.QUEST-SFT-Data-Open-ended
QUEST SFT Data (Open-ended)
Project Page | Paper | GitHub
Open-ended supervised fine-tuning trajectories for QUEST (tool-using assistant format). Split: train. Columns: messages (list[{role, content}]).
Load
from datasets import load_dataset
ds = load_dataset("osunlp/QUEST-SFT-Data-Open-ended", split="train", streaming=True)
row = next(iter(ds))
print(row.keys())
QUEST Family
Type
Resources
35B checkpoints
RL, MT+SFT, MT, SFT
30B checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/QUEST-SFT-Data-Open-ended.open-christian-data
Open Christian Data
Open Christian Data (OCD) aims to be the single unified collection of all public domain Christian text. It exists to bring this collection together from across the internet and to structure it in a useful format for public use.
Beyond the Bible, Christian writing is poorly represented as a cohesive dataset or as data structured for AI training. This Hugging Face release is the AI-focused publication of the collection: consistent, downloadable JSON for model… See the full description on the dataset page: https://huggingface.co/datasets/OpenChristianDataOrg/open-christian-data.ODA-Fin-RL-12k
Unlocking Data Value in Finance: A Study on Distillation
and Difficulty-Aware Training
📖 Overview
ODA-Fin-RL-12K is a carefully curated dataset for reinforcement learning (RL) in financial domain, comprising 12,187 hard-but-verifiable samples. Designed to complement ODA-Fin-SFT-318K, this dataset targets challenging financial reasoning tasks with concise, reliably verifiable answers—optimized for RL training.
🎯 Key Highlights
12K Hard Samples: Curated… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Fin-RL-12k.Open_Reaction_Data
ORDerly: Styrene Mizoroki-Heck RAG-Ready Dataset
This repository contains chemical reaction data formatted for Retrieval-Augmented Generation (RAG) systems.
The data is a processed version of the ORDerly benchmark, specifically focusing on reaction conditions and forward/retro prediction tasks.
Dataset Structure
The data is split into 10,000-row Parquet chunks to prevent Out-of-Memory (OOM) errors during ingestion into vector databases.
It includes:
orderly_condition:… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Open_Reaction_Data.SlimPajama-Meta-rater-Cleanliness-30B
Top 30B token SlimPajama Subset selected by the Cleanliness rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Cleanliness dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Cleanliness-30B.SA-Prot-annot
SA-Prot-Annot Dataset (Sci-Align)
🌌 The Sciverse Data Foundation
Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research.
Sciverse consists of three core data… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SA-Prot-annot.K12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究
morocco-cassation-court-decisions
Morocco Cassation Court Decisions
29,000+ full-text decisions from the Moroccan Court of Cassation (محكمة النقض)Source: juriscassation.cspj.ma — Official portal of the Supreme Council of the Judiciary (CSPJ)License: CC BY 4.0
Why this dataset exists
In 2026, accessing the jurisprudence of the Court of Cassation in Morocco requires being physically located in Morocco and armed with patience. The official website does not allow searching by date range, imposes a… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataMoroccanLaw/morocco-cassation-court-decisions.gohumanize-open-humanizer-dataset
GoHumanize Open Humanizer Dataset
2,957 training pairs and 300 test pairs for teaching a language model to rewrite
AI-styled English prose into natural human writing. Each pair is:
input: a passage rewritten by a large language model in the register typical of LLM output
(formal, smooth, hedged, connective phrases, no contractions);
output: the original human-written passage, from a public-domain book or, since version 2,
from a US federal government publication.
The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.Snowball-67B-A2B-RLVR1-Repro-Data
Snowball 67B-A2B RLVR1 data
These are the exact Parquet inputs retained for the Snowball 67B-A2B sync and
async RLVR1 experiments on Iris cw-rno2a in September 2026. The data was
selected from the skyrl_gym route of a TaskTrove conversion of the public
NVIDIA Nemotron RL Ultra training blend,
preserving source order and holding out the last 100 selected rows. See
provenance.json for the local conversion and filtering record. The original
TaskTrove release
is also public.… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-RLVR1-Repro-Data.thai-open-data-go-th
Thai Open Data.GO.TH
This dataset collects all text from data.go.th (CC-BY licence Only) and is pre-processing the data to the pre-training model.
List Datasets:
folktoys: https://data.go.th/dataset/folktoys
folktales: https://data.go.th/dataset/folktales
ARDA: https://data.go.th/dataset/blog
Thai Rituals: https://data.go.th/dataset/rituals
Thai old community area: https://data.go.th/th/dataset/gdpublish-http-ocd-onep-go-th
Thai aqualib: https://data.go.th/th/dataset/aqualib
Thai… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-open-data-go-th.pedro-open-dataset-1kitalian-open-sft-chat-dataset
Italian Open SFT Chat Dataset
An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records.
This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.pedro-open-dataset-max-512-tokenspedro-open-dataset-max-512-tokens-10k
