datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-pool-7b-2xdatacomp_pools
DataComp Pools
This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.dolma3_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3 pool, pre–quality upsampling and mixing.
If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025.
Dolma 3 Pool
The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.dolma3.5_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3.5 pool. It contains no quality upsampling or mixing. This is an updated version of the Dolma 3 pool with additional quality filtering and more data sources.
If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025.
Dolma 3.5 Pool
The Dolma 3.5 pool is a dataset of nearly 10 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3.5_pool.dclm-pool-1b-1xdolma3_dolmino_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3 Dolmino pool; it hasn't been mixed.
If you are interested in the data used to train:
Olmo 3 7B: allenai/dolma3_dolmino_mix-100B-1025
Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125
Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training
This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 7B.
Dataset Sources
Source
Category
Tokens
Documents
TinyMATH Mind… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_pool.dcvlm_pool_large
DCVLM-Pool (large)
The raw candidate pool at the large scale of our DataComp-VLM
benchmark: 1,949,321,868 samples / 166.7 TB across 166 source datasets, as
WebDataset tar shards — ≈4× the medium pool.
🚚 Upload in progress
This repo is being populated incrementally and is not yet complete — shards are still being
uploaded. Sources already present are final and safe to use; sources with fewer shards than the counts
quoted below have not finished uploading yet.… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_large.dclm-pool-7b-1xdcvlm_pool_medium
DCVLM-Pool (medium)
The raw candidate pool at the medium scale of our DataComp-VLM
benchmark: 483,576,747 samples / 41.1 TB across 166 source datasets, as
WebDataset tar shards — ≈4× the small pool.
This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose
the filters and the mixing ratios, and create another training set. If you instead want a
ready-to-train dataset, use dcvlm-baseline-200b
(our reference SoTA… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_medium.dolma3_longmino_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3 Longmino pool; it hasn't been mixed.
If you are interested in the data used to train:
Olmo 3 7B: allenai/dolma3_longmino_mix-50B-1025
Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125
Dolma 3 Longmino Pool (639B)
Dolma 3 Longmino Pool is the full pool of documents considered for stage 3 (long context) extension trainin of Olmo 3 7B.
Dataset Sources
Source
Type
Tokens
Docs
LC-s2pdf-REX 32k-64k
Synth PDFs
24.1B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_longmino_pool.dclm-pool-1b-5xdcvlm_pool_small
DCVLM-Pool (small)
The raw candidate pool at the small scale of our DataComp-VLM
benchmark: 120,940,134 samples / ~187.5B tokens / 10.3 TB across 166 source datasets, as
WebDataset tar shards.
This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose
the filters and the mixing ratios, and create another training set. If you instead want a
ready-to-train dataset, use dcvlm-baseline-200b
(our reference SoTA DCVLM-baseline… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_small.Raon-OpenTTS-Pool
Raon-OpenTTS-Pool
Technical Report
Raon-OpenTTS-Pool is a large-scale open English speech corpus for text-to-speech (TTS) training,
constructed from 8 publicly available speech corpora and a set of web-sourced recordings.
It is the training data behind Raon-OpenTTS,
an open TTS model that performs on par with state-of-the-art closed-data systems.
615K hours of speech audio
239.7M speech segments
11 source datasets aggregated into a unified format
All… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Pool.cascade-eval-pool
cascade eval pool — lagged public reveal (exact bytes)
Retired snapshots of the held-out evaluation pool used by the
cascade subnet. Each folder is a
byte-identical mirror of the pool/snapshots/block-<N>.tar that validators
scored — downloaded from the private pool bucket, sha256-verified against the
publisher index, and republished unmodified. A snapshot is revealed only after
a newer snapshot has superseded it, so no revealed pool can be selected by a
current or future round.… See the full description on the dataset page: https://huggingface.co/datasets/Tensor-Link/cascade-eval-pool.locus-commit-pool-v1
Locus Commit Pool v1
Native Git history, preserved as replayable software changes
Commit message · complete selected before-state · unified patches · native object IDs · provenance · experimental labels
Locus Commit Pool v1 is a large evidence pool for studying and training on how real software changes. Each document represents one surviving single-parent, multi-file Git commit. It keeps the commit message, the selected files as they existed before the… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-commit-pool-v1.igris-poolCPT_Data_Pool
CPT Data Pool
This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training.
For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo.
Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.dclm-pool-400m-1xarchive-dolma3-pool-150b-enriched
archive-dolma3-pool-150b-enriched
ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B pool sample.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_150B_enriched
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory including… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b-enriched.patchrecoverygym-laguna
PatchRecoveryGym for Laguna
Submitted by: Kannappan Sirchabesan (@kannappans) · Poolside Research Hackathon (Foundations track)
A reproducible eval + RL environment that tests whether a coding agent can
recover from a wrong first attempt — a real, under-measured agentic-coding
weakness. Built for Poolside Laguna XS.2 on dependency-migration repair tasks.
📦 Installable Verifiers environment on the Prime Hub · 🎯 deterministic hidden-test reward · 🔁 144-candidate reranking… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/patchrecoverygym-laguna.archive-dolma3-pool-150b
archive-dolma3-pool-150b
ARCHIVE (pre-6T era): the 150B-token Dolma3 pool sample. Kept for reproducibility of earlier work; current attribution work uses the 6T datasets.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_150B
Renamed
2026-05-25
See… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b.tmax-image-pool
tmax image pool
Every Apptainer/SIF image referenced by the three swerl-tmax-15k variants, stored once
and shared between them.
Read the layout from the repo, not from a prefix
The plain images are split across four directories, not one. A consumer that filters
on images/ silently gets 4,519 fewer files and then reports itself complete — the
tasks that reference the missing blobs fail later, at sandbox start, one by one.
directory
files
images/
9,971… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/tmax-image-pool.ethereum-attestation-pool
Ethereum attestation pool
Slot-level observations comparing attestations visible in a consensus client's pending pool with attestations subsequently included in blocks. The panel supports analysis of local pool coverage and inclusion counts.
Contents
Table
Record
ethereum_attestation_observations
Attester-slot counts observed pending and included, with their difference and collection coverage
Using the data
attesters_seen_in_pool… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/ethereum-attestation-pool.bitcoin-mining-pool-templates
Bitcoin mining pool templates
Timestamped Stratum job messages collected directly from Bitcoin mining pool endpoints. The data records changes in the work each endpoint sends to miners, including the previous block hash, coinbase data and clean-jobs flag.
Contents
Table
Record
bitcoin_mining_pool_jobs
A job received from a pool endpoint, with its observation time, nTime, coinbase, merkle branch count and clean-jobs flag
Using the data… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-mining-pool-templates.SPADE-Environment-Pool-GPT5.5-Games
SPARE GPT-5.5 Grounded Cognitive Multi-Turn Games
This public dataset contains 7,872 validated Python game environments for actor-only SPARE training.
Six cognitive skills, exactly 1,312 environments per skill
Generated with GPT-5.5 and grounded by spice_megascience_15k.jsonl
Grounding corpus SHA-256: a36a928b4940b5b5d9e3f4cb5804a94c69462360943adb3be14613c82f0f72c0
Maximum 25 turns and 32K generation context
Every environment passes load, reset, step, and replay validation with… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-Games.watercolour-reference-pool
Watercolour reference pool
The reference paintings that define the reward in the watercolour RL environment: an
agent writes a p5.brush sketch, the sketch is
rendered, and a vision judge compares the render against paintings sampled from this pool.
What the pool contains is the reward function. Replace it and you have changed what
the environment rewards, without touching a line of code.
178 paintings in two tiers, each with the JavaScript source that produced it.
tier… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-reference-pool.InsPLAD-workshop-pool
Dataset Card for InsPLAD Workshop Pool
This is a FiftyOne dataset with 1754 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/InsPLAD-workshop-pool")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/InsPLAD-workshop-pool.dcvlm_pool_small_annotations
DCVLM-Pool (small) — per-sample annotations
Every filtering annotation we computed for the small data pool of our
DataComp-VLM paper: image quality, image–text alignment, language ID,
text-quality classifiers, multimodal perplexity, decontamination scores and more — up to 167 fields per
sample (180 distinct fields overall), for all 120,940,134 samples across 166 source datasets.
These are the raw annotations, not a filtered dataset. They are the inputs our curation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_small_annotations.locus-repo-code-pool-v2
Locus repository code pool
The Locus repository code pool contains curated repository-version documents
for code-language-model research. Each row combines the useful files from one
Git repository snapshot into a single deterministic text document while
retaining source provenance, file order, language and test signals, health
evidence, and stable content hashes.
This release is a candidate acquisition pool, not a ready-made training
split. Use a globally deduplicated manifest… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-repo-code-pool-v2.datacomp-medium-pool-translated
