CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /dolma3_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 pool, pre–quality upsampling and mixing. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3 Pool The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.texttext-generation10B<n<100B41 likes28k downloads7mo agoHugging Face02allenai /dolma3_mix-6T-1025-7B ⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️ For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T. Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B. For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.texttext-generation1B<n<10B56 likes15k downloads8mo agoHugging Face03allenai /dolma3_dolmino_mix-100B-1025 Dolma 3 Dolmino Mix (100B) The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model. Dataset Sources Source Category Tokens Documents TinyMATH Mind Math (synth) 898M (0.9%) 1.52M TinyMATH PoT Math (synth) 241M (0.24%) 758K CraneMath Math (synth) 5.62B (5.63%) 7.24M MegaMatt Math (synth) 1.73B (1.73%) 3.23M Dolmino Math Math (synth) 10.7B (10.7%) 22.3M StackEdu (FIM) Code 10.0B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1025.texttext-generation10M<n<100M10 likes12k downloads9mo agoHugging Face04allenai /dolma3_mix-150B-1025 Dolma 3 Sample: 150B Mix Dataset Sources Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3 Source Type Tokens Documents Common Crawl Web pages 121B (76.9%) 84.5M olmOCR Science PDFs Academic documents 19.9B (12.6%) 2.25M Stack-Edu (Rebalanced) GitHub code 11.1B (7.06%) 14.3M arXiv Papers with LaTeX 1.29B (0.82%) 247K FineMath 3+ Math web pages 4.10B (2.60%) 2.57M Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.texttext-generation10M<n<100M10 likes8k downloads8mo agoHugging Face05salmankhanpm /dolma3_dolmino_mix-100B-1025 Dolma 3 Dolmino Mix (100B) The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model. Dataset Sources Source Category Tokens Documents TinyMATH Mind Math (synth) 898M (0.9%) 1.52M TinyMATH PoT Math (synth) 241M (0.24%) 758K CraneMath Math (synth) 5.62B (5.63%) 7.24M MegaMatt Math (synth) 1.73B (1.73%) 3.23M Dolmino Math Math (synth) 10.7B (10.7%) 22.3M StackEdu (FIM) Code 10.0B… See the full description on the dataset page: https://huggingface.co/datasets/salmankhanpm/dolma3_dolmino_mix-100B-1025.texttext-generation10M<n<100M0 likes5.2k downloads8mo agoHugging Face06HCAI-Lab-GT /dolma3-6t-sample-10000-docs dolma3-6t-sample-10000-docs Materialized stratified sample of 10K docs per bin (5.68M total docs, 10.5B tokens). Seed 42. This is the basis for the SOC-156 TrackStar gradient index. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_6T_sample_10000_docs Renamed 2026-05-25… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs.text1M<n<10M0 likes2.5k downloads4mo agoHugging Face07TheFinAI /dolma3_300B_sampletabular100M<n<1B0 likes1.9k downloads4mo agoHugging Face08orionweller /dolma_20bn_wiki_upsampletabular10M<n<100M0 likes1.1k downloads2y agoHugging Face09orionweller /dolma_20bn_instruct_upsampletabular10M<n<100M0 likes718 downloads2y agoHugging Face10skymizer /dolma-v1_7-50Btexttext-generation10M<n<100M0 likes704 downloads2y agoHugging Face11placeholderlabs /exp-pool-repository-code-dolma2-tokenized Locus EXP Repository Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes664 downloads1mo agoHugging Face12skymizer /dolma-v1_7-50B-1-of-2text10M<n<100M0 likes654 downloads2y agoHugging Face13allenai /dolma3_dolmino_mix-10B-1025 Dolma 3 dolmino dataset mix (10B) for Olmo3 stage 2 annealing training. Smaller mixture of high-quality data. 10B is the size at which we ran all micro anneals for Olmo 3 stage 2 training. For the full mixture, please see: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1025 Licensing Information Dolma 3 Dolmino is licensed under the Open Data Commons Attribution License v1.0 (ODC-By). It is intended for research and educational use. For more… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-10B-1025.texttext-generation1M<n<10M10 likes613 downloads9mo agoHugging Face14nhagar /dolma_urls_v1.6 Dataset Card for dolma_urls_v1.6 This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.6.text1B<n<10B0 likes565 downloads1y agoHugging Face15orionweller /dolma_20bn_cc_high_qualitytabular10M<n<100M0 likes513 downloads2y agoHugging Face16Hugodonotexit /dolma-en Dolma-English Overview This dataset is a filtered subset of the Dolma corpus, restricted to English-only documents and further constrained by a minimum document length threshold. It is intended for training and evaluating large language models and other NLP systems that benefit from higher-quality, sufficiently long English text. The primary goals of this dataset are: To reduce multilingual and very short/noisy content present in the raw Dolma corpus. To provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/dolma-en.texttext-generation1B<n<10B1 likes503 downloads7mo agoHugging Face17emozilla /dolma-v1_7-arxivtext1M<n<10M3 likes502 downloads2y agoHugging Face18orionweller /dolma_18bn_stratified_sampletabular10M<n<100M0 likes499 downloads2y agoHugging Face19emozilla /dolma-v1_7-cc_en_headtext100M<n<1B1 likes492 downloads2y agoHugging Face20orionweller /dolma_20bn_prop_stratified_sampletabular10M<n<100M0 likes485 downloads2y agoHugging Face21wonabru /dolma-books-chunked-4ktext1M<n<10M1 likes484 downloads1y agoHugging Face22emozilla /dolma-v1_7-305BThis dataset is a 10% sample of Dolma v1.7, equating to around ~305B tokens and uploaded directly as a Hugging Face dataset. As a pure sample, it maintains the ODC-BY license. texttext-generation100M<n<1B11 likes420 downloads2y agoHugging Face23emozilla /dolma-v1_7-c4text100M<n<1B2 likes409 downloads2y agoHugging Face24orionweller /dolma_20bn_no_instructtabular10M<n<100M0 likes409 downloads2y agoHugging Face25nkandpa2 /wiki-dolma Wiki Datasets Preprocessed versions of openly licensed wiki dumps collected by wikiteam and hosted on the Internet Archive. Version Descriptions raw: The original wikitext v0: Wikitext parsed to plain text with wtf\_wikipedia and conversion of math templates to LaTeX. v1: Removal of some html snippets left behind during parsing. v2: Removal of documents that basically just transcripts of non-openly licensed things. v3: Removal of documents that basically… See the full description on the dataset page: https://huggingface.co/datasets/nkandpa2/wiki-dolma.text10M<n<100M0 likes374 downloads2y agoHugging Face26orionweller /dolma_18bn_prop_stratified_sampletabular10M<n<100M0 likes362 downloads2y agoHugging Face27emozilla /dolma-v1_7-30BThis dataset is a 1% sample of Dolma v1.7, equating to around ~30B tokens and uploaded directly as a Hugging Face dataset. As a pure sample, it maintains the ODC-BY license. texttext-generation10M<n<100M2 likes361 downloads2y agoHugging Face28nhagar /dolma_urls_v1.5 Dataset Card for dolma_urls_v1.5 This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.5.text1B<n<10B0 likes343 downloads1y agoHugging Face29nhagar /dolma_urls_v1.7 Dataset Card for dolma_urls_v1.7 This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.7.text1B<n<10B0 likes326 downloads1y agoHugging Face30ericrcwu /week1-general-20b-dolma2-v1 Week-One General 20B Dolma2 This is a deterministic, pretokenized 20-billion-token baseline corpus for controlled language-model architecture and training experiments. It contains nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered extension of the previous view. It also includes a dataset-only 370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count. The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.texttext-generationn<1K0 likes317 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.