CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /dolma3_mix-6T Dolma 3 Mix (6T) The Dolma 3 Mix (6T) is the collection of data used during the pretraining stage to train the Olmo-3-1125-32B model. This dataset is made up of ~6 trillion tokens from a diverse mix of web content, academic publications, code, and more. The majority of this dataset comes from Common Crawl. For more information on Dolma, please see our original release here. Smaller Sample for Analysis Available! If you would like a smaller sample of this mix… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T.text-generation36 likes78k downloads8mo agoHugging Face02allenai /dolma3.5_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3.5 pool. It contains no quality upsampling or mixing. This is an updated version of the Dolma 3 pool with additional quality filtering and more data sources. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3.5 Pool The Dolma 3.5 pool is a dataset of nearly 10 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3.5_pool.text-generation8 likes71k downloads3mo agoHugging Face03allenai /dolma3_dolmino_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 Dolmino pool; it hasn't been mixed. If you are interested in the data used to train: Olmo 3 7B: allenai/dolma3_dolmino_mix-100B-1025 Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125 Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 7B. Dataset Sources Source Category Tokens Documents TinyMATH Mind… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_pool.text-generation8 likes37k downloads9mo agoHugging Face04allenai /dolma3_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 pool, pre–quality upsampling and mixing. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3 Pool The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.texttext-generation10B<n<100B41 likes28k downloads7mo agoHugging Face05allenai /dolma3_dolmino_mix-100B-1125 Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 32B. Dataset Sources Source Category TinyMATH Mind Math (synth) TinyMATH PoT Math (synth) CraneMath Math (synth) MegaMatt Math (synth) Dolmino Math Math (synth) StackEdu (FIM) Code CraneCode Python (synth) Reddit To Flashcards QA (synth) Wiki To RCQA QA (synth) Nemotron Synth QA QA… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1125.25 likes21k downloads7mo agoHugging Face06allenai /dolma3_mix-6T-1025-7B ⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️ For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T. Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B. For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.texttext-generation1B<n<10B56 likes15k downloads8mo agoHugging Face07allenai /dolma3_longmino_mix-100B-1125 Dolma 3 Longmino Mix (100B) The Dolma 3 Longmino Mix (100B) is the mixture of data used for the third stage of training for Olmo 3 32B model. Dataset Sources Source Type LC-s2pdf-REX 32k-64k Synth PDFs LC-s2pdf-CWE 32k-64k Synth PDFs LC-s2pdf 32k-64k PDFs LC-s2pdf 8k-32k (8-16k) PDFs LC-s2pdf 8k-32k (16-32k) PDFs Midtraining Data Mix Licensing Information Dolma 3 Longmino is licensed under the Open Data Commons Attribution… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_longmino_mix-100B-1125.18 likes13k downloads7mo agoHugging Face08allenai /dolma3_longmino_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 Longmino pool; it hasn't been mixed. If you are interested in the data used to train: Olmo 3 7B: allenai/dolma3_longmino_mix-50B-1025 Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125 Dolma 3 Longmino Pool (639B) Dolma 3 Longmino Pool is the full pool of documents considered for stage 3 (long context) extension trainin of Olmo 3 7B. Dataset Sources Source Type Tokens Docs LC-s2pdf-REX 32k-64k Synth PDFs 24.1B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_longmino_pool.14 likes13k downloads9mo agoHugging Face09allenai /dolma3_dolmino_mix-100B-1025 Dolma 3 Dolmino Mix (100B) The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model. Dataset Sources Source Category Tokens Documents TinyMATH Mind Math (synth) 898M (0.9%) 1.52M TinyMATH PoT Math (synth) 241M (0.24%) 758K CraneMath Math (synth) 5.62B (5.63%) 7.24M MegaMatt Math (synth) 1.73B (1.73%) 3.23M Dolmino Math Math (synth) 10.7B (10.7%) 22.3M StackEdu (FIM) Code 10.0B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1025.texttext-generation10M<n<100M10 likes12k downloads9mo agoHugging Face10allenai /dolma3_mix-150B-1025 Dolma 3 Sample: 150B Mix Dataset Sources Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3 Source Type Tokens Documents Common Crawl Web pages 121B (76.9%) 84.5M olmOCR Science PDFs Academic documents 19.9B (12.6%) 2.25M Stack-Edu (Rebalanced) GitHub code 11.1B (7.06%) 14.3M arXiv Papers with LaTeX 1.29B (0.82%) 247K FineMath 3+ Math web pages 4.10B (2.60%) 2.57M Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.texttext-generation10M<n<100M10 likes8k downloads8mo agoHugging Face11allenai /dolma3_longmino_mix-50B-1025 Dolma 3 Longmino Mix (50B) The Dolma 3 Longmino Mix (50B) is the mixture of data used for the third stage of training for Olmo 3 7B model. Dataset Sources Source Type Tokens Docs LC-s2pdf-REX 32k-64k Synth PDFs 6.08B (12.2%) 217K LC-s2pdf-CWE 32k-64k Synth PDFs 1.94B (3.88%) 71.3K LC-s2pdf 32k-64k PDFs 4.81B (9.63%) 177K LC-s2pdf 8k-32k (8-16k) PDFs 2.27B (4.55%) 235K LC-s2pdf 8k-32k (16-32k) PDFs 1.85B (3.70%) 110K Midtraining Data Mix 33.0B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_longmino_mix-50B-1025.10 likes7.4k downloads9mo agoHugging Face12salmankhanpm /dolma3_dolmino_mix-100B-1025 Dolma 3 Dolmino Mix (100B) The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model. Dataset Sources Source Category Tokens Documents TinyMATH Mind Math (synth) 898M (0.9%) 1.52M TinyMATH PoT Math (synth) 241M (0.24%) 758K CraneMath Math (synth) 5.62B (5.63%) 7.24M MegaMatt Math (synth) 1.73B (1.73%) 3.23M Dolmino Math Math (synth) 10.7B (10.7%) 22.3M StackEdu (FIM) Code 10.0B… See the full description on the dataset page: https://huggingface.co/datasets/salmankhanpm/dolma3_dolmino_mix-100B-1025.texttext-generation10M<n<100M0 likes5.2k downloads8mo agoHugging Face13HCAI-Lab-GT /dolma3-6t-sample-100000-docs dolma3-6t-sample-100000-docs Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/). Layout HCAI-Lab/dolma3-6t-sample-100000-docs/ ├── bin_summary.csv ├── sample_contract.json └── worker_NNNN/ └── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker) Dual-access… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-100000-docs.0 likes5.1k downloads4mo agoHugging Face14HCAI-Lab-GT /archive-dolma3-mix-150b-enriched archive-dolma3-mix-150b-enriched ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B Dolma3 mix (not the pool). Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_mix_150B_enriched Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-mix-150b-enriched.0 likes4.4k downloads4mo agoHugging Face15HCAI-Lab-GT /archive-dolma3-pool-150b-enriched archive-dolma3-pool-150b-enriched ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B pool sample. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_pool_150B_enriched Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory including… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b-enriched.0 likes4.1k downloads4mo agoHugging Face16allenai /dolmaDolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Researchtext-generationn>1T1.1k likes3.9k downloads2y agoHugging Face17HCAI-Lab-GT /archive-dolma3-pool-150b archive-dolma3-pool-150b ARCHIVE (pre-6T era): the 150B-token Dolma3 pool sample. Kept for reproducibility of earlier work; current attribution work uses the 6T datasets. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_pool_150B Renamed 2026-05-25 See… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b.1 likes3.4k downloads4mo agoHugging Face18pico-lm /pretokenized-dolma The Pretokenized Dolma Dataset A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library. Overview Key Features: Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280 Sequence length: 2049 tokens (2048 + 1 for next-token prediction) Sharded into 10,000 Parquet files (~78MB each) 420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.100M<n<1B4 likes3k downloads1y agoHugging Face19HCAI-Lab-GT /dolma3-6t-sample-5000-docs dolma3-6t-sample-5000-docs Materialized stratified sample of 5K docs per bin (2.86M total docs, 5.3B tokens). Seed 42. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_5000_docs Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-5000-docs.0 likes2.5k downloads4mo agoHugging Face20HCAI-Lab-GT /dolma3-6t-sample-10000-docs dolma3-6t-sample-10000-docs Materialized stratified sample of 10K docs per bin (5.68M total docs, 10.5B tokens). Seed 42. This is the basis for the SOC-156 TrackStar gradient index. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_10000_docs Renamed 2026-05-25… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs.text1M<n<10M0 likes2.5k downloads4mo agoHugging Face21TheFinAI /dolma3_300B_sampletabular100M<n<1B0 likes1.9k downloads4mo agoHugging Face22HCAI-Lab-GT /archive-dolma3-6t-dedup-state archive-dolma3-6t-dedup-state ARCHIVE: backup of the SOC-90 dedup state (Bloom filter + filtered docs). Reproducibility-only; current dedup state is in HCAI-Lab/dolma3-6t-bloom-index and dolma3-6t-unique. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/soc-90-keep-backup… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-6t-dedup-state.tabularn<1K0 likes1.6k downloads4mo agoHugging Face23yangwang92 /dolma3-2T0 likes1.5k downloads1mo agoHugging Face24HCAI-Lab-GT /dolma3-6t-corpus-manifest dolma3-6t-corpus-manifest Unified per-document manifest joining topic/format/quality/token-count/source-shard for the Dolma3 6T corpus. Cross-shard parquet partitioned dataset. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_corpus_manifest Renamed 2026-05-25 See… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-corpus-manifest.0 likes1.4k downloads4mo agoHugging Face25HCAI-Lab-GT /dolma3-6t-unique dolma3-6t-unique 1.258B unique Dolma3 documents materialized after Bloom-filter dedup (SOC-90). Used as the canonical deduplicated 6T corpus on HF. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6t_unique Renamed 2026-05-25 See docs/data_home/inventory.json for the… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-unique.1 likes1.4k downloads4mo agoHugging Face26HCAI-Lab-GT /dolma3-6t-sample-1000-docs dolma3-6t-sample-1000-docs Materialized stratified sample of 1K docs per bin (575K total docs, 1.08B tokens). Seed 42. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_1000_docs Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-1000-docs.0 likes1.3k downloads4mo agoHugging Face27orionweller /dolma_20bn_wiki_upsampletabular10M<n<100M0 likes1.1k downloads2y agoHugging Face28HCAI-Lab-GT /dolma3-6t-preconditioner-100k dolma3-6t-preconditioner-100k 100K uniform-random preconditioner sample (251M tokens, 38K unique source shards). For TrackStar / EK-FAC preconditioner construction. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_preconditioner_100k Renamed 2026-05-25 See… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-preconditioner-100k.1 likes1k downloads4mo agoHugging Face29emozilla /dolma-v1_7-305B-tokenized-llama2-nanosetimagen<1K0 likes906 downloads2y agoHugging Face30HCAI-Lab-GT /dolma3-6t-sample-500-docs dolma3-6t-sample-500-docs Materialized stratified sample of 500 docs per bin (288K total docs, 539M tokens) drawn from the deduplicated 6T Dolma3 corpus. Seed 42. Matching companion bucket at hf://buckets/HCAI-Lab/dolma3-6t-sample-500-docs. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-500-docs.0 likes867 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.