datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ling-Coder-SFT
🤗 Hugging Face
🤖 ModelScope
🖥️ GitHub
Ling-Coder Dataset
The Ling-Coder Dataset comprises the following components:
Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples.
Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples.
Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SFT.forum-competition-math-training-pool
Forum competition mathematics training pool
Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and
shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file
format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the
union of those same datasets in one format, one JSON object per line, deduplicated by problem text
and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.AReaL-tau2-data
AReaL-tau2-data
Synthetic training data for multi-turn interactive tool-using agents, generated by SEA, a self-evolving multi-agent data engine. This dataset is used to train AReaL-SEA-235B-A22B, achieving state-of-the-art results on τ²-bench.
Paper: From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
Training Framework: AReaL
Benchmark: τ²-bench
Dataset Overview
The dataset covers three customer-service… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/AReaL-tau2-data.olympiad-math-training-pool
Olympiad mathematics training pool
Public olympiad and competition mathematics, four datasets gathered at pinned revisions, shipped
twice over. sources/ holds each dataset the way its publisher ships it, in its own file format
with its own fields and nothing renamed, 229052 rows across four folders. pool/ holds the union
of those same datasets in one format, one JSON object per line, deduplicated by problem text and
reduced to 225822 rows, every row labelled with the dataset it… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/olympiad-math-training-pool.SWE-CARE
SWE-CARE: A Comprehensiveness-aware Benchmark for Code Review Evaluation
Dataset Description
SWE-CARE (Software Engineering - Comprehensive Analysis and Review Evaluation) is a comprehensiveness-aware benchmark for evaluating Large Language Models (LLMs) on repository-level code review tasks. The dataset features real-world code review scenarios from popular open-source Python and Java repositories, with comprehensive metadata and… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/SWE-CARE.incantation-elden-ring-scenes
Incantation Elden Ring Combat Captions
Paper | Project page | GitHub
Preview subset. This repository is a public preview and reference subset of the Incantation dataset. It documents the data format, annotation style, and initial training material used by the project. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper.
This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/zhush/incantation-elden-ring-scenes.PubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.Ling-Coder-SyntheticQA
🤗 Hugging Face
🤖 ModelScope
🖥️ GitHub
Ling-Coder Dataset
The Ling-Coder Dataset comprises the following components:
Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples.
Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples.
Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SyntheticQA.competition-math-training-pool
Competition mathematics training pool
Public competition mathematics, six datasets gathered at pinned revisions, shipped twice over.
sources/ holds each dataset the way its publisher ships it, in its own file format with its own
fields and nothing renamed, 1951046 rows across six folders. pool/ holds the union of those
same datasets in one format, one JSON object per line, deduplicated by problem text and reduced
to 1125451 rows, every row labelled with the dataset it came from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-math-training-pool.joyo-kanji-yomi-benchmark-parakeet
日本語 | English
常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet)
常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。
このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。
このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。
概要… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/joyo-kanji-yomi-benchmark-parakeet.PubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/Bilsteen/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.A3S-Bench
Agent3Sigma-Stage (A3S-Bench)
💻 GitHub | 🏆 Leaderboard | 📄 Paper (PDF) | arXiv
Agent3Sigma-Stage (A3S-Bench) is an end-to-end security evaluation framework for autonomous agents (e.g., OpenClaw), designed to systematically measure both an Agent's ability to resist attacks during multi-turn interactions and its utility in completing legitimate tasks. The framework provides an evaluation dataset covering 10 security risk categories across 6 real-world usage scenarios… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/A3S-Bench.LiteResearcher-SFT-Data
LiteResearcher — SFT Cold-Start Data
Distilled deep-research trajectories used for the SFT cold-start of LiteResearcher-4B
This dataset contains the 68,231 multi-turn deep-research trajectories used to
train the SFT cold-start checkpoint that RL (GRPO+TIS) is later launched from —
the "68.2 K distilled deep-research trajectories" referenced in the paper and in
LiteResearcher-Data.
Each row is a complete ReAct-style episode: a research question, the model's
interleaved… See the full description on the dataset page: https://huggingface.co/datasets/simplex-ai-inc/LiteResearcher-SFT-Data.task070_abductivenli_incorrect_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task070_abductivenli_incorrect_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task070_abductivenli_incorrect_classification.Ring-lite-sft-data
🤖 ModelScope
🤗 HuggingFace
🖥️ GitHub
Ring-lite-sft-data
This is a the SFT data used during the fine-tuning of the Ring-lite model. The query pool was sourced from open-source repositories and further enriched through synthetic generation using large language models (LLMs). To ensure the production of high-fidelity responses with Long-CoT, we implemented an iterative refinement pipeline that synergistically combines automated model… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ring-lite-sft-data.task454_swag_incorrect_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task454_swag_incorrect_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task454_swag_incorrect_answer_generation.logic-grid-puzzles-training-pool
Logic grid puzzles training pool
Logic grid puzzles: a row of positions, a handful of attributes with one value per position, and a
list of clues that together admit exactly one arrangement. Two sets drawn for this pool by
generators run here under the seeds recorded below, and two public datasets read at the pinned
revisions named below, laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 390945 rows, one JSON… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/logic-grid-puzzles-training-pool.task068_abductivenli_incorrect_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task068_abductivenli_incorrect_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task068_abductivenli_incorrect_answer_generation.python-unit-test-training-pool
Python unit test training pool
A pool of public data for training a model to write tests for Python code. It is a
straight collection of open datasets, not a new corpus: every row comes from one of the
sources below, at the revision named, and the only rows removed are the ones an overlap
filter flagged against held-out material this pool is kept separate from.
Every row of the normalised layer pairs a program with tests for it. That is the point of
the pool, and it is why the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-unit-test-training-pool.PubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/th-laurel/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.code-execution-trace-training-pool
Code execution trace training pool
Public Python code paired with one concrete call and the value that call returns. Every value in
this pool was computed by running the code, not copied from a label. The data is laid out twice,
and either layer may be used.
pool/
Every source rewritten into one shape, 2174322 rows over 11 gzipped parts, one JSON object per
line, with these fields.
Field
What it holds
id
a row identifier unique within this pool
code… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/code-execution-trace-training-pool.symbolic-calculus-training-pool
Symbolic calculus training pool
Single variable calculus exercises with closed form answers: derivatives of composite expressions,
indefinite and definite integrals, limits of indeterminate forms, and coefficients of Maclaurin
series, with some multivariable operators in one of the sources. A set generated for this pool and
two public datasets read at the pinned revisions named below, laid out twice. Train on either
layer or on both.
pool.jsonl
Every source… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/symbolic-calculus-training-pool.constrained-instruction-training-pool
Constrained instruction training pool
Public prompts for writing tasks, many of them carrying a constraint a program can check, from ten
datasets read at the pinned revisions named below and one layer built here from them. The pool is
laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 462652 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/constrained-instruction-training-pool.incantation-elden-ring-scenes
Incantation Elden Ring Combat Captions
Preview subset. This repository is an early public preview and reference subset of the Incantation dataset. It is provided to document the data format, annotation style, and initial training material ahead of the full paper/project release. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper.
This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/MatrixTeam/incantation-elden-ring-scenes.task043_essential_terms_answering_incomplete_questions
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task043_essential_terms_answering_incomplete_questions
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task043_essential_terms_answering_incomplete_questions.long-context-retrieval-training-pool
Long context retrieval training pool
Long prompts with short, checkable answers. Each row is one complete message: a task instruction,
a long body of text that hides what the question is about, and the question itself, together with
every string an answer has to contain for it to be right. The bodies run from four thousand to
thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on
both.
pool.jsonl
Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.in-context-grid-reasoning
In-Context Grid Reasoning (ICGR)
A small, fully synthetic benchmark for demonstration-conditioned rule induction:
each task shows 2–4 (input grid → output grid) support pairs that share one
hidden transformation, and the model must apply the same transformation to a
held-out query input.
It targets the same behaviour probed by recent in-context / latent-reasoning work
on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning,
arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.task297_storycloze_incorrect_end_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task297_storycloze_incorrect_end_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task297_storycloze_incorrect_end_classification.python-functions-training-pool
Python function-writing training pool
A pool of public data for training a model to write Python functions. It is a straight
collection of open datasets, not a new corpus: every row comes from one of the sources
below, at the revision named, and the only rows removed are the ones an overlap filter
flagged against held-out material this pool is kept separate from.
Rows in the normalised layer: 5756045.
Rows in the raw layer: 6258415.
The two layers
pool/ holds the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-functions-training-pool.Ring-lite-rl-data
🤗 Hugging Face
🤖 ModelScope
🖥️ GitHub
Ring-lite-rl-data
This dataset is a curated subset of high-quality problems across mathematics and code domains designed for reinforcement learning in the Ring-lite model. This dataset contains:
Mathematics: Over 39,000 rigorously curated problems sourced from:
Open-source datasets (BigMath, DeepScaleR, DAPO, DeepMath-103K)
Art of Problem Solving (AoPS) contest collections
Code: Approximately 8,400… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ring-lite-rl-data.
