datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolma3_pool⚠️ IMPORTANT NOTICE ⚠️
This is the Dolma 3 pool, pre–quality upsampling and mixing.
If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025.
Dolma 3 Pool
The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.dolma3_mix-6T-1025-7B
⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️
For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T.
Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B.
For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.dolma3_dolmino_mix-100B-1025
Dolma 3 Dolmino Mix (100B)
The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model.
Dataset Sources
Source
Category
Tokens
Documents
TinyMATH Mind
Math (synth)
898M (0.9%)
1.52M
TinyMATH PoT
Math (synth)
241M (0.24%)
758K
CraneMath
Math (synth)
5.62B (5.63%)
7.24M
MegaMatt
Math (synth)
1.73B (1.73%)
3.23M
Dolmino Math
Math (synth)
10.7B (10.7%)
22.3M
StackEdu (FIM)
Code
10.0B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1025.dolma3_mix-150B-1025
Dolma 3 Sample: 150B Mix
Dataset Sources
Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3
Source
Type
Tokens
Documents
Common Crawl
Web pages
121B (76.9%)
84.5M
olmOCR Science PDFs
Academic documents
19.9B (12.6%)
2.25M
Stack-Edu (Rebalanced)
GitHub code
11.1B (7.06%)
14.3M
arXiv
Papers with LaTeX
1.29B (0.82%)
247K
FineMath 3+
Math web pages
4.10B (2.60%)
2.57M
Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.dolma3_dolmino_mix-100B-1025
Dolma 3 Dolmino Mix (100B)
The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model.
Dataset Sources
Source
Category
Tokens
Documents
TinyMATH Mind
Math (synth)
898M (0.9%)
1.52M
TinyMATH PoT
Math (synth)
241M (0.24%)
758K
CraneMath
Math (synth)
5.62B (5.63%)
7.24M
MegaMatt
Math (synth)
1.73B (1.73%)
3.23M
Dolmino Math
Math (synth)
10.7B (10.7%)
22.3M
StackEdu (FIM)
Code
10.0B… See the full description on the dataset page: https://huggingface.co/datasets/salmankhanpm/dolma3_dolmino_mix-100B-1025.dolma3-6t-sample-10000-docs
dolma3-6t-sample-10000-docs
Materialized stratified sample of 10K docs per bin (5.68M total docs, 10.5B tokens). Seed 42. This is the basis for the SOC-156 TrackStar gradient index.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_10000_docs
Renamed
2026-05-25… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs.dolma3_300B_sampledolma_20bn_wiki_upsampledolma_20bn_instruct_upsampledolma-v1_7-50Bexp-pool-repository-code-dolma2-tokenized
Locus EXP Repository Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.dolma-v1_7-50B-1-of-2dolma3_dolmino_mix-10B-1025
Dolma 3 dolmino dataset mix (10B) for Olmo3 stage 2 annealing training.
Smaller mixture of high-quality data. 10B is the size at which we ran all micro anneals for Olmo 3 stage 2 training. For the full mixture, please see: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1025
Licensing Information
Dolma 3 Dolmino is licensed under the Open Data Commons Attribution License v1.0 (ODC-By). It is intended for research and educational use. For more… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-10B-1025.dolma_urls_v1.6
Dataset Card for dolma_urls_v1.6
This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.6.dolma_20bn_cc_high_qualitydolma-en
Dolma-English
Overview
This dataset is a filtered subset of the Dolma corpus, restricted to English-only documents and further constrained by a minimum document length threshold. It is intended for training and evaluating large language models and other NLP systems that benefit from higher-quality, sufficiently long English text.
The primary goals of this dataset are:
To reduce multilingual and very short/noisy content present in the raw Dolma corpus.
To provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/dolma-en.dolma-v1_7-arxivdolma_18bn_stratified_sampledolma-v1_7-cc_en_headdolma_20bn_prop_stratified_sampledolma-books-chunked-4kdolma-v1_7-305BThis dataset is a 10% sample of Dolma v1.7, equating to around ~305B tokens and uploaded directly as a Hugging Face dataset.
As a pure sample, it maintains the ODC-BY license.
dolma-v1_7-c4dolma_20bn_no_instructwiki-dolma
Wiki Datasets
Preprocessed versions of openly licensed wiki dumps collected by wikiteam and hosted on the Internet Archive.
Version Descriptions
raw: The original wikitext
v0: Wikitext parsed to plain text with wtf\_wikipedia and conversion of math templates to LaTeX.
v1: Removal of some html snippets left behind during parsing.
v2: Removal of documents that basically just transcripts of non-openly licensed things.
v3: Removal of documents that basically… See the full description on the dataset page: https://huggingface.co/datasets/nkandpa2/wiki-dolma.dolma_18bn_prop_stratified_sampledolma-v1_7-30BThis dataset is a 1% sample of Dolma v1.7, equating to around ~30B tokens and uploaded directly as a Hugging Face dataset.
As a pure sample, it maintains the ODC-BY license.
dolma_urls_v1.5
Dataset Card for dolma_urls_v1.5
This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.5.dolma_urls_v1.7
Dataset Card for dolma_urls_v1.7
This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.7.week1-general-20b-dolma2-v1
Week-One General 20B Dolma2
This is a deterministic, pretokenized 20-billion-token baseline corpus for
controlled language-model architecture and training experiments. It contains
nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered
extension of the previous view. It also includes a dataset-only
370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B
order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.
