datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolma3_mix-150B-1025
Dolma 3 Sample: 150B Mix
Dataset Sources
Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3
Source
Type
Tokens
Documents
Common Crawl
Web pages
121B (76.9%)
84.5M
olmOCR Science PDFs
Academic documents
19.9B (12.6%)
2.25M
Stack-Edu (Rebalanced)
GitHub code
11.1B (7.06%)
14.3M
arXiv
Papers with LaTeX
1.29B (0.82%)
247K
FineMath 3+
Math web pages
4.10B (2.60%)
2.57M
Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.wiki-dolma
Wiki Datasets
Preprocessed versions of openly licensed wiki dumps collected by wikiteam and hosted on the Internet Archive.
Version Descriptions
raw: The original wikitext
v0: Wikitext parsed to plain text with wtf\_wikipedia and conversion of math templates to LaTeX.
v1: Removal of some html snippets left behind during parsing.
v2: Removal of documents that basically just transcripts of non-openly licensed things.
v3: Removal of documents that basically… See the full description on the dataset page: https://huggingface.co/datasets/nkandpa2/wiki-dolma.week1-general-20b-dolma2-v1
Week-One General 20B Dolma2
This is a deterministic, pretokenized 20-billion-token baseline corpus for
controlled language-model architecture and training experiments. It contains
nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered
extension of the previous view. It also includes a dataset-only
370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B
order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.dolma3-6t-sample-10000-docs-finance-and-business
HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business
Filename-derived finance_and_business slice of
HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to
revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7.
Extraction rule
The corpus contains every source .jsonl.zst file whose filename contains
the literal segment -finance_and_business-. Source paths and compressed file
contents are preserved byte-for-byte. This is a coarse WebOrganizer
finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.dolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.dolma-reddit-to-flashcards-0625
Overview
Dolma Reddit to Flashcards is a dataset of synthetically-generated QA items created on the basis of filtered Reddit data.
The creation of this dataset was motivated by the observation in Dolma (Soldaini et al. 2024) that the original Dolma Reddit data showed no benefit from inclusion of thread-level context over isolated submissions and comments, and that clean performance distinctions between tested Reddit versions were limited mainly to the HellaSWAG benchmark.
The… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma-reddit-to-flashcards-0625.allenai-dolma3_mix-6T-samplearchive-dolma3-pool-stratified
archive-dolma3-pool-stratified
ARCHIVE (pre-6T era): stratified sampling outputs from the 150B pool work.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_stratified
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory including the… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-stratified.dolma-v1_7-50B-second-phase
Dataset Overview
This dataset is based on the original AllenAI Dolma dataset and contains the estimated split of sources used for the second phase of training OLMo.
For additional context on the dataset methodology and its impact, please refer to the blog post: OLMO 1.7 7B: A 24-Point Improvement on MMLU.
The dataset metadata allows users to dynamically select which parts to download — saving both storage and bandwidth.
Original Split:
Name
Source
Total Tokens… See the full description on the dataset page: https://huggingface.co/datasets/luca-g97/dolma-v1_7-50B-second-phase.ai-culture-multilingual-json-dolma
AI-Culture Multilingual JSON + DOLMA Corpus
16M words · 12 languages · CC-BY-4.0
The AI-Culture corpus contains 5K articles providing comprehensive philosophical and cultural content, exploring the intersection of technology, artificial intelligence, and human culture, perfectly aligned across 12 languages. All content maintains identical parallel structure across translations with zero duplication and editor-curated quality.
This project is maintained by a non-profit digital… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/ai-culture-multilingual-json-dolma.arxiv_dolmastackexchange_dolmaarchive-dolma3-enriched-small
archive-dolma3-enriched-small
ARCHIVE (pre-6T era): small enriched Dolma sample.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma_enriched_small
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory including the old_names field on each entry.
dolma_low_quality_datamediawiki-dolma
MediaWiki Datasets
dolma3_pool_staging⚠️ TESTING ONLY - DO NOT USE ⚠️
This is a staging repository for testing internal Dolma 3 processing pipeline. It contains no useful data. If you are looking for the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T and allenai/dolma3_pool.
dolma-miniDPI-Dolmadolma_web1_raw
Raw data for Thai Dolma
Commoncrawl
Fineweb2
thai-dolma-web_v9backup-thai_dolma-gamble_webbackup-thai_dolma-gamble_web_fullthai-dolma-web_v11_our_ccthai-dolma-web_v10
