datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolma3_mix-6T
Dolma 3 Mix (6T)
The Dolma 3 Mix (6T) is the collection of data used during the pretraining stage to train the Olmo-3-1125-32B model. This dataset is made up of ~6 trillion tokens from a diverse mix of web content, academic publications, code, and more. The majority of this dataset comes from Common Crawl.
For more information on Dolma, please see our original release here.
Smaller Sample for Analysis Available!
If you would like a smaller sample of this mix… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T.dolma3.5_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3.5 pool. It contains no quality upsampling or mixing. This is an updated version of the Dolma 3 pool with additional quality filtering and more data sources.
If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025.
Dolma 3.5 Pool
The Dolma 3.5 pool is a dataset of nearly 10 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3.5_pool.dolma3_dolmino_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3 Dolmino pool; it hasn't been mixed.
If you are interested in the data used to train:
Olmo 3 7B: allenai/dolma3_dolmino_mix-100B-1025
Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125
Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training
This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 7B.
Dataset Sources
Source
Category
Tokens
Documents
TinyMATH Mind… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_pool.dolma3_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3 pool, pre–quality upsampling and mixing.
If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025.
Dolma 3 Pool
The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.dolma3_dolmino_mix-100B-1125
Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training
This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 32B.
Dataset Sources
Source
Category
TinyMATH Mind
Math (synth)
TinyMATH PoT
Math (synth)
CraneMath
Math (synth)
MegaMatt
Math (synth)
Dolmino Math
Math (synth)
StackEdu (FIM)
Code
CraneCode
Python (synth)
Reddit To Flashcards
QA (synth)
Wiki To RCQA
QA (synth)
Nemotron Synth QA
QA… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1125.dolma3_mix-6T-1025-7B
⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️
For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T.
Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B.
For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.dolma3_longmino_mix-100B-1125
Dolma 3 Longmino Mix (100B)
The Dolma 3 Longmino Mix (100B) is the mixture of data used for the third stage of training for Olmo 3 32B model.
Dataset Sources
Source
Type
LC-s2pdf-REX 32k-64k
Synth PDFs
LC-s2pdf-CWE 32k-64k
Synth PDFs
LC-s2pdf 32k-64k
PDFs
LC-s2pdf 8k-32k (8-16k)
PDFs
LC-s2pdf 8k-32k (16-32k)
PDFs
Midtraining Data
Mix
Licensing Information
Dolma 3 Longmino is licensed under the Open Data Commons Attribution… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_longmino_mix-100B-1125.dolma3_longmino_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3 Longmino pool; it hasn't been mixed.
If you are interested in the data used to train:
Olmo 3 7B: allenai/dolma3_longmino_mix-50B-1025
Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125
Dolma 3 Longmino Pool (639B)
Dolma 3 Longmino Pool is the full pool of documents considered for stage 3 (long context) extension trainin of Olmo 3 7B.
Dataset Sources
Source
Type
Tokens
Docs
LC-s2pdf-REX 32k-64k
Synth PDFs
24.1B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_longmino_pool.dolma3_dolmino_mix-100B-1025
Dolma 3 Dolmino Mix (100B)
The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model.
Dataset Sources
Source
Category
Tokens
Documents
TinyMATH Mind
Math (synth)
898M (0.9%)
1.52M
TinyMATH PoT
Math (synth)
241M (0.24%)
758K
CraneMath
Math (synth)
5.62B (5.63%)
7.24M
MegaMatt
Math (synth)
1.73B (1.73%)
3.23M
Dolmino Math
Math (synth)
10.7B (10.7%)
22.3M
StackEdu (FIM)
Code
10.0B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1025.dolma3_mix-150B-1025
Dolma 3 Sample: 150B Mix
Dataset Sources
Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3
Source
Type
Tokens
Documents
Common Crawl
Web pages
121B (76.9%)
84.5M
olmOCR Science PDFs
Academic documents
19.9B (12.6%)
2.25M
Stack-Edu (Rebalanced)
GitHub code
11.1B (7.06%)
14.3M
arXiv
Papers with LaTeX
1.29B (0.82%)
247K
FineMath 3+
Math web pages
4.10B (2.60%)
2.57M
Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.dolma3_longmino_mix-50B-1025
Dolma 3 Longmino Mix (50B)
The Dolma 3 Longmino Mix (50B) is the mixture of data used for the third stage of training for Olmo 3 7B model.
Dataset Sources
Source
Type
Tokens
Docs
LC-s2pdf-REX 32k-64k
Synth PDFs
6.08B (12.2%)
217K
LC-s2pdf-CWE 32k-64k
Synth PDFs
1.94B (3.88%)
71.3K
LC-s2pdf 32k-64k
PDFs
4.81B (9.63%)
177K
LC-s2pdf 8k-32k (8-16k)
PDFs
2.27B (4.55%)
235K
LC-s2pdf 8k-32k (16-32k)
PDFs
1.85B (3.70%)
110K
Midtraining Data
Mix
33.0B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_longmino_mix-50B-1025.dolma3_dolmino_mix-100B-1025
Dolma 3 Dolmino Mix (100B)
The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model.
Dataset Sources
Source
Category
Tokens
Documents
TinyMATH Mind
Math (synth)
898M (0.9%)
1.52M
TinyMATH PoT
Math (synth)
241M (0.24%)
758K
CraneMath
Math (synth)
5.62B (5.63%)
7.24M
MegaMatt
Math (synth)
1.73B (1.73%)
3.23M
Dolmino Math
Math (synth)
10.7B (10.7%)
22.3M
StackEdu (FIM)
Code
10.0B… See the full description on the dataset page: https://huggingface.co/datasets/salmankhanpm/dolma3_dolmino_mix-100B-1025.dolma3-6t-sample-100000-docs
dolma3-6t-sample-100000-docs
Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/).
Layout
HCAI-Lab/dolma3-6t-sample-100000-docs/
├── bin_summary.csv
├── sample_contract.json
└── worker_NNNN/
└── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker)
Dual-access… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-100000-docs.archive-dolma3-mix-150b-enriched
archive-dolma3-mix-150b-enriched
ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B Dolma3 mix (not the pool).
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_mix_150B_enriched
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-mix-150b-enriched.archive-dolma3-pool-150b-enriched
archive-dolma3-pool-150b-enriched
ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B pool sample.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_150B_enriched
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory including… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b-enriched.dolmaDolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Researcharchive-dolma3-pool-150b
archive-dolma3-pool-150b
ARCHIVE (pre-6T era): the 150B-token Dolma3 pool sample. Kept for reproducibility of earlier work; current attribution work uses the 6T datasets.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_150B
Renamed
2026-05-25
See… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b.pretokenized-dolma
The Pretokenized Dolma Dataset
A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library.
Overview
Key Features:
Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280
Sequence length: 2049 tokens (2048 + 1 for next-token prediction)
Sharded into 10,000 Parquet files (~78MB each)
420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.dolma3-6t-sample-5000-docs
dolma3-6t-sample-5000-docs
Materialized stratified sample of 5K docs per bin (2.86M total docs, 5.3B tokens). Seed 42.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_5000_docs
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-5000-docs.dolma3-6t-sample-10000-docs
dolma3-6t-sample-10000-docs
Materialized stratified sample of 10K docs per bin (5.68M total docs, 10.5B tokens). Seed 42. This is the basis for the SOC-156 TrackStar gradient index.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_10000_docs
Renamed
2026-05-25… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs.dolma3_300B_samplearchive-dolma3-6t-dedup-state
archive-dolma3-6t-dedup-state
ARCHIVE: backup of the SOC-90 dedup state (Bloom filter + filtered docs). Reproducibility-only; current dedup state is in HCAI-Lab/dolma3-6t-bloom-index and dolma3-6t-unique.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/soc-90-keep-backup… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-6t-dedup-state.dolma3-2Tdolma3-6t-corpus-manifest
dolma3-6t-corpus-manifest
Unified per-document manifest joining topic/format/quality/token-count/source-shard for the Dolma3 6T corpus. Cross-shard parquet partitioned dataset.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_corpus_manifest
Renamed
2026-05-25
See… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-corpus-manifest.dolma3-6t-unique
dolma3-6t-unique
1.258B unique Dolma3 documents materialized after Bloom-filter dedup (SOC-90). Used as the canonical deduplicated 6T corpus on HF.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6t_unique
Renamed
2026-05-25
See docs/data_home/inventory.json for the… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-unique.dolma3-6t-sample-1000-docs
dolma3-6t-sample-1000-docs
Materialized stratified sample of 1K docs per bin (575K total docs, 1.08B tokens). Seed 42.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_1000_docs
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-1000-docs.dolma_20bn_wiki_upsampledolma3-6t-preconditioner-100k
dolma3-6t-preconditioner-100k
100K uniform-random preconditioner sample (251M tokens, 38K unique source shards). For TrackStar / EK-FAC preconditioner construction.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_preconditioner_100k
Renamed
2026-05-25
See… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-preconditioner-100k.dolma-v1_7-305B-tokenized-llama2-nanosetdolma3-6t-sample-500-docs
dolma3-6t-sample-500-docs
Materialized stratified sample of 500 docs per bin (288K total docs, 539M tokens) drawn from the deduplicated 6T Dolma3 corpus. Seed 42. Matching companion bucket at hf://buckets/HCAI-Lab/dolma3-6t-sample-500-docs.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-500-docs.
