CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /dolma3_mix-150B-1025 Dolma 3 Sample: 150B Mix Dataset Sources Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3 Source Type Tokens Documents Common Crawl Web pages 121B (76.9%) 84.5M olmOCR Science PDFs Academic documents 19.9B (12.6%) 2.25M Stack-Edu (Rebalanced) GitHub code 11.1B (7.06%) 14.3M arXiv Papers with LaTeX 1.29B (0.82%) 247K FineMath 3+ Math web pages 4.10B (2.60%) 2.57M Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.texttext-generation10M<n<100M10 likes8k downloads8mo agoHugging Face02nkandpa2 /wiki-dolma Wiki Datasets Preprocessed versions of openly licensed wiki dumps collected by wikiteam and hosted on the Internet Archive. Version Descriptions raw: The original wikitext v0: Wikitext parsed to plain text with wtf\_wikipedia and conversion of math templates to LaTeX. v1: Removal of some html snippets left behind during parsing. v2: Removal of documents that basically just transcripts of non-openly licensed things. v3: Removal of documents that basically… See the full description on the dataset page: https://huggingface.co/datasets/nkandpa2/wiki-dolma.text10M<n<100M0 likes374 downloads2y agoHugging Face03ericrcwu /week1-general-20b-dolma2-v1 Week-One General 20B Dolma2 This is a deterministic, pretokenized 20-billion-token baseline corpus for controlled language-model architecture and training experiments. It contains nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered extension of the previous view. It also includes a dataset-only 370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count. The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.texttext-generationn<1K0 likes319 downloads2mo agoHugging Face04HCAI-Lab-GT /dolma3-6t-sample-10000-docs-finance-and-business HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business Filename-derived finance_and_business slice of HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7. Extraction rule The corpus contains every source .jsonl.zst file whose filename contains the literal segment -finance_and_business-. Source paths and compressed file contents are preserved byte-for-byte. This is a coarse WebOrganizer finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.texttext-generation100K<n<1M0 likes281 downloads1mo agoHugging Face05Mr-Philo /dolma3_dolmino_megatron_tokenize Dolma 3 / Dolmino Megatron-LM indexed dataset This repository contains immutable Megatron-LM indexed datasets (.bin and .idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally contains no training checkpoints, experiment outputs, logs, or dataset caches. The indexed payloads were derived from these pinned public datasets: allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.tabulartext-generationn<1K0 likes194 downloads25d agoHugging Face06allenai /dolma-reddit-to-flashcards-0625 Overview Dolma Reddit to Flashcards is a dataset of synthetically-generated QA items created on the basis of filtered Reddit data. The creation of this dataset was motivated by the observation in Dolma (Soldaini et al. 2024) that the original Dolma Reddit data showed no benefit from inclusion of thread-level context over isolated submissions and comments, and that clean performance distinctions between tested Reddit versions were limited mainly to the HellaSWAG benchmark. The… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma-reddit-to-flashcards-0625.text100M<n<1B7 likes167 downloads1y agoHugging Face07agentlans /allenai-dolma3_mix-6T-sampletexttext-generation100K<n<1M0 likes146 downloads3mo agoHugging Face08HCAI-Lab-GT /archive-dolma3-pool-stratified archive-dolma3-pool-stratified ARCHIVE (pre-6T era): stratified sampling outputs from the 150B pool work. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_pool_stratified Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory including the… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-stratified.tabular10M<n<100M0 likes114 downloads4mo agoHugging Face09luca-g97 /dolma-v1_7-50B-second-phase Dataset Overview This dataset is based on the original AllenAI Dolma dataset and contains the estimated split of sources used for the second phase of training OLMo. For additional context on the dataset methodology and its impact, please refer to the blog post: OLMO 1.7 7B: A 24-Point Improvement on MMLU. The dataset metadata allows users to dynamically select which parts to download — saving both storage and bandwidth. Original Split: Name Source Total Tokens… See the full description on the dataset page: https://huggingface.co/datasets/luca-g97/dolma-v1_7-50B-second-phase.tabular100M<n<1B0 likes113 downloads1y agoHugging Face10AI-Culture-Commons /ai-culture-multilingual-json-dolma AI-Culture Multilingual JSON + DOLMA Corpus 16M words · 12 languages · CC-BY-4.0 The AI-Culture corpus contains 5K articles providing comprehensive philosophical and cultural content, exploring the intersection of technology, artificial intelligence, and human culture, perfectly aligned across 12 languages. All content maintains identical parallel structure across translations with zero duplication and editor-curated quality. This project is maintained by a non-profit digital… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/ai-culture-multilingual-json-dolma.texttranslation1K<n<10K3 likes92 downloads1y agoHugging Face11orionweller /arxiv_dolmatext1M<n<10M0 likes48 downloads2y agoHugging Face12orionweller /stackexchange_dolmatext10M<n<100M2 likes34 downloads2y agoHugging Face13HCAI-Lab-GT /archive-dolma3-enriched-small archive-dolma3-enriched-small ARCHIVE (pre-6T era): small enriched Dolma sample. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma_enriched_small Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory including the old_names field on each entry. text10K<n<100K0 likes22 downloads4mo agoHugging Face14Elain-q /dolma_low_quality_datatext10K<n<100K0 likes12 downloads3y agoHugging Face15nkandpa2 /mediawiki-dolma MediaWiki Datasets text10M<n<100M0 likes11 downloads2y agoHugging Face16allenai /dolma3_pool_staging⚠️ TESTING ONLY - DO NOT USE ⚠️ This is a staging repository for testing internal Dolma 3 processing pipeline. It contains no useful data. If you are looking for the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T and allenai/dolma3_pool. texttext-generationn<1K1 likes11 downloads7mo agoHugging Face17wannaphong /dolma-minitext10M<n<100M0 likes8 downloads11mo agoHugging Face18DataProvenanceInitiative /DPI-Dolmatext1M<n<10M1 likes7 downloads2y agoHugging Face19wannaphong /dolma_web1_rawgated Raw data for Thai Dolma Commoncrawl Fineweb2 texttext-generation100M<n<1B0 likes2 downloads1y agoHugging Face20wannaphong /thai-dolma-web_v9gatedtext100M<n<1B0 likes1 downloads1y agoHugging Face21wannaphong /backup-thai_dolma-gamble_webgatedimage10M<n<100M0 likes1 downloads1y agoHugging Face22wannaphong /backup-thai_dolma-gamble_web_fullgatedimage100M<n<1B0 likes1 downloads1y agoHugging Face23wannaphong /thai-dolma-web_v11_our_ccgatedtext10M<n<100M0 likes1 downloads1y agoHugging Face24wannaphong /thai-dolma-web_v10gatedtext100M<n<1B0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.