datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolma3.5_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3.5 pool. It contains no quality upsampling or mixing. This is an updated version of the Dolma 3 pool with additional quality filtering and more data sources.
If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025.
Dolma 3.5 Pool
The Dolma 3.5 pool is a dataset of nearly 10 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3.5_pool.dolma3_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3 pool, pre–quality upsampling and mixing.
If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025.
Dolma 3 Pool
The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.dolma3_dolmino_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3 Dolmino pool; it hasn't been mixed.
If you are interested in the data used to train:
Olmo 3 7B: allenai/dolma3_dolmino_mix-100B-1025
Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125
Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training
This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 7B.
Dataset Sources
Source
Category
Tokens
Documents
TinyMATH Mind… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_pool.locus-commit-pool-v1
Locus Commit Pool v1
Native Git history, preserved as replayable software changes
Commit message · complete selected before-state · unified patches · native object IDs · provenance · experimental labels
Locus Commit Pool v1 is a large evidence pool for studying and training on how real software changes. Each document represents one surviving single-parent, multi-file Git commit. It keeps the commit message, the selected files as they existed before the… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-commit-pool-v1.patchrecoverygym-laguna
PatchRecoveryGym for Laguna
Submitted by: Kannappan Sirchabesan (@kannappans) · Poolside Research Hackathon (Foundations track)
A reproducible eval + RL environment that tests whether a coding agent can
recover from a wrong first attempt — a real, under-measured agentic-coding
weakness. Built for Poolside Laguna XS.2 on dependency-migration repair tasks.
📦 Installable Verifiers environment on the Prime Hub · 🎯 deterministic hidden-test reward · 🔁 144-candidate reranking… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/patchrecoverygym-laguna.exp-pool-repository-code-dolma2-tokenized
Locus EXP Repository Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.locus-repo-code-pool-v2
Locus repository code pool
The Locus repository code pool contains curated repository-version documents
for code-language-model research. Each row combines the useful files from one
Git repository snapshot into a single deterministic text document while
retaining source provenance, file order, language and test signals, health
evidence, and stable content hashes.
This release is a candidate acquisition pool, not a ready-made training
split. Use a globally deduplicated manifest… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-repo-code-pool-v2.forum-competition-math-training-pool
Forum competition mathematics training pool
Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and
shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file
format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the
union of those same datasets in one format, one JSON object per line, deduplicated by problem text
and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.olympiad-math-training-pool
Olympiad mathematics training pool
Public olympiad and competition mathematics, four datasets gathered at pinned revisions, shipped
twice over. sources/ holds each dataset the way its publisher ships it, in its own file format
with its own fields and nothing renamed, 229052 rows across four folders. pool/ holds the union
of those same datasets in one format, one JSON object per line, deduplicated by problem text and
reduced to 225822 rows, every row labelled with the dataset it… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/olympiad-math-training-pool.dart-math-pool-gsm8k
[!NOTE]
This dataset is the data pool synthesized from the query set of the GSM8K training set,
containing all answer-correct samples and other metadata produced during the work.
DART-Math-* datasets are extracted from dart-math-pool-* data pools.
🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub
🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX
Datasets:… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-gsm8k.competition-math-training-pool
Competition mathematics training pool
Public competition mathematics, six datasets gathered at pinned revisions, shipped twice over.
sources/ holds each dataset the way its publisher ships it, in its own file format with its own
fields and nothing renamed, 1951046 rows across six folders. pool/ holds the union of those
same datasets in one format, one JSON object per line, deduplicated by problem text and reduced
to 1125451 rows, every row labelled with the dataset it came from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-math-training-pool.exp-pool-commit-code-dolma2-tokenized
Locus EXP Commit Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.dart-math-pool-math
[!NOTE]
This dataset is the data pool synthesized from the query set of the MATH training set,
containing all answer-correct samples and other metadata produced during the work.
DART-Math-* datasets are extracted from dart-math-pool-* data pools.
🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub
🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX
Datasets:… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-math.exp-pool-academic-dolma2-tokenized
Locus EXP Academic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.exp-pool-olmo-web-dolma2-tokenized
Locus EXP OLMo Web - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.pool-qkbue
传奇私服分布式路由与自动化接口索引库 - Batch 002
本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。
📂 区域节点集群子目录 (Spider Pool Indexes)
👉 今日传奇私服网 - 传奇私服推荐 - 新开传奇私服 —— 承载源站 🌐 alpha.agergames.com
👉 传奇私服推荐 - 新开传奇私服 - 传奇私服发布 —— 承载源站 🌐 alpha.amoygame.com
👉 传奇私服999 - 传奇私服 - 传奇私服网盘 —— 承载源站 🌐 alpha.gamecctv.com
👉 传奇私服 - 传奇私服推荐网 - 传奇私服999 —— 承载源站 🌐 alpha.gamerling.com
👉 传奇私服发布 - 传奇私服网 - 今日传奇私服 —— 承载源站 🌐 alpha.games-vr.com
👉 新开传奇私服发布 - 传奇私服网 - 传奇私服 —— 承载源站 🌐 alpha.j5games.com
👉 传奇私服 -… See the full description on the dataset page: https://huggingface.co/datasets/keiseoq/pool-qkbue.exp-pool-encyclopedic-dolma2-tokenized
Locus EXP Encyclopedic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.exp-pool-nemotron-math-dolma2-tokenized
Locus EXP Nemotron Math - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.pool-g4cl
WhatsApp网页版分布式路由与自动化接口索引库 - Batch 005
本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:WhatsApp网页版)。
📂 区域节点集群子目录 (Spider Pool Indexes)
👉 WhatsApp网页版-质量纠纷处理-WhatsApp Web 网页版 —— 承载源站 🌐 Agile.alliance-whatapp.hl.cn
👉 WhatsApp网页版-未读消息一键过滤-WhatsApp Web 网页版 —— 承载源站 🌐 Agile.alphas-whatapp.hl.cn
👉 WhatsApp网页版-WhatsApp Web 网页版-CRM侧边栏同步 —— 承载源站 🌐 Agile.badge-whatapp.hl.cn
👉 WhatsApp网页版-全局搜索历史凭证-WhatsApp Web 网页版 —— 承载源站 🌐 Agile.battle-whatapp.hl.cn
👉 WhatsApp网页版-WhatsApp Web… See the full description on the dataset page: https://huggingface.co/datasets/makesi1/pool-g4cl.logic-grid-puzzles-training-pool
Logic grid puzzles training pool
Logic grid puzzles: a row of positions, a handful of attributes with one value per position, and a
list of clues that together admit exactly one arrangement. Two sets drawn for this pool by
generators run here under the seeds recorded below, and two public datasets read at the pinned
revisions named below, laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 390945 rows, one JSON… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/logic-grid-puzzles-training-pool.hs3-prompt-pool-topic-judged
hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training
Prompts only (no completions). Every user prompt in
model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated
35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0)
for the high-level topic of both quirk families.
Why
Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.python-unit-test-training-pool
Python unit test training pool
A pool of public data for training a model to write tests for Python code. It is a
straight collection of open datasets, not a new corpus: every row comes from one of the
sources below, at the revision named, and the only rows removed are the ones an overlap
filter flagged against held-out material this pool is kept separate from.
Every row of the normalised layer pairs a program with tests for it. That is the point of
the pool, and it is why the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-unit-test-training-pool.code-execution-trace-training-pool
Code execution trace training pool
Public Python code paired with one concrete call and the value that call returns. Every value in
this pool was computed by running the code, not copied from a label. The data is laid out twice,
and either layer may be used.
pool/
Every source rewritten into one shape, 2174322 rows over 11 gzipped parts, one JSON object per
line, with these fields.
Field
What it holds
id
a row identifier unique within this pool
code… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/code-execution-trace-training-pool.symbolic-calculus-training-pool
Symbolic calculus training pool
Single variable calculus exercises with closed form answers: derivatives of composite expressions,
indefinite and definite integrals, limits of indeterminate forms, and coefficients of Maclaurin
series, with some multivariable operators in one of the sources. A set generated for this pool and
two public datasets read at the pinned revisions named below, laid out twice. Train on either
layer or on both.
pool.jsonl
Every source… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/symbolic-calculus-training-pool.constrained-instruction-training-pool
Constrained instruction training pool
Public prompts for writing tasks, many of them carrying a constraint a program can check, from ten
datasets read at the pinned revisions named below and one layer built here from them. The pool is
laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 462652 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/constrained-instruction-training-pool.long-context-retrieval-training-pool
Long context retrieval training pool
Long prompts with short, checkable answers. Each row is one complete message: a task instruction,
a long body of text that hides what the question is about, and the question itself, together with
every string an answer has to contain for it to be right. The bodies run from four thousand to
thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on
both.
pool.jsonl
Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.pool-encyclopedic
locus-1 encyclopedic pool
The encyclopedic pool for locus-1 - one of seven cluster
datasets that are mixed into the pretraining corpus. Every pool shares one schema, so a
mixture is a query rather than a rebuild.
6,498,683 documents, 7,314,613,330 tokens under allenai/dolma2-tokenizer@5292e5d6c0f40b67cc765fe41bec991cf4345b5c
Built from: HuggingFaceFW/finewiki (en @ 8bd13e72e6a0)
Build settings digest: cbd87eae472c6b30
Fingerprint compatibility digest: 6ee2a61c2f91d6a1… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pool-encyclopedic.python-functions-training-pool
Python function-writing training pool
A pool of public data for training a model to write Python functions. It is a straight
collection of open datasets, not a new corpus: every row comes from one of the sources
below, at the revision named, and the only rows removed are the ones an overlap filter
flagged against held-out material this pool is kept separate from.
Rows in the normalised layer: 5756045.
Rows in the raw layer: 6258415.
The two layers
pool/ holds the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-functions-training-pool.procedural-reasoning-training-pool
Procedural reasoning training pool
Reasoning questions from 101 procedural generators, each of which writes a question, computes its
own answer and ships a verifier that scores an attempt at it, plus a collection of solved Sudoku
puzzles. Every answer is short and exactly checkable, so a trained model can be marked against the
key by a program and no judge is needed. Laid out twice. Train on either layer or on both.
pool.jsonl
Every generator rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/procedural-reasoning-training-pool.competition-answer-math-training-pool
Competition answer mathematics training pool
Competition mathematics problems that ask for a single final answer, three public datasets gathered
at pinned revisions, shipped twice over. sources/ holds each dataset the way its publisher ships
it, in its own file format with its own fields and nothing renamed, 113045 rows across three
folders. pool/ holds the union of those same datasets in one format, one JSON object per line,
deduplicated by problem text and reduced to 107637… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-answer-math-training-pool.
