CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /dclm-baseline-1.0 DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.tabular1B<n<10B315 likes635k downloads2y agoHugging Face02mlfoundations /dclm-pool-7b-2x3 likes155k downloads2y agoHugging Face03mlfoundations /dclm-pool-1b-1x3 likes33k downloads2y agoHugging Face04mlfoundations /dclm-baseline-1.0-parquet DCLM-baseline Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format. DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.tabular1B<n<10B56 likes20k downloads2y agoHugging Face05EssentialAI /eai-taxonomy-code-w-dclm 💻 EAI-Taxonomy Code w/ DCLM 🏆 Website | 🖥️ Code | 📖 Paper A 564 billion token dataset of high-quality code curated from web data using taxonomy-based filtering. 🎯 Dataset Overview This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional code datasets that require complex domain-specific pipelines, our approach leverages a 12-category taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-code-w-dclm.texttext-generation100M<n<1B10 likes18k downloads1y agoHugging Face06mlfoundations /dclm-pool-7b-1x1 likes17k downloads2y agoHugging Face07mlfoundations /dclm-pool-1b-5x1 likes12k downloads2y agoHugging Face08Zyphra /dclm-dedup DCLM-Deduped DCLM is a recently released high quality dataset that uses model-based quality filtering to filter a large subset of common-crawl for similarity to OpenHermes and other instruction-tuning datasets. For reference see the DCLM paper. The original authors of DCLM did not release fully deduplicated version of their dataset, claiming that full deduplication did not improve performance. The released version was partially deduplicated in shards. Nevertheless, when performing… See the full description on the dataset page: https://huggingface.co/datasets/Zyphra/dclm-dedup.tabulartext-generation100M<n<1B22 likes5k downloads2y agoHugging Face09mlfoundations /dclm-pool-400m-1x3 likes4.4k downloads2y agoHugging Face10gair-prox /DCLM-pro 📚 DCLM-pro ArXiv | Models | Code DCLM-pro is refined from DCLM using the ProX refining framework. It contains about >500B high quality tokens, ready for general language model pre-training. License DCLM-pro is based on DCLM, which is made available under an cc-by-4.0 license. Citation @article{zhou2024programming, title={Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale}, author={Zhou, Fan and Wang, Zengzhi… See the full description on the dataset page: https://huggingface.co/datasets/gair-prox/DCLM-pro.texttext-generation100M<n<1B13 likes4.3k downloads2y agoHugging Face11semran1 /dclm-stem-filteredtext1M<n<10M0 likes3.3k downloads1y agoHugging Face12EleutherAI /dclm-dedup_20250227-004105tabular100M<n<1B1 likes3.2k downloads2y agoHugging Face13HuggingFaceFW /finepdfs_50BT-dclm_30BT-fineweb_edu_20BT FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing. Dataset Description Component Source Tokens FinePDFs 100BT FinePDFs ~50B DCLM 100BT DCLM-Baseline 1.0 ~30B FineWeb-Edu 100BT FineWeb-Edu ~20B The schema is reduced to the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.text10M<n<100M2 likes2.8k downloads7mo agoHugging Face14HuggingFaceTB /dclm-edu DCLM-Edu Description This is a filtered version of DCLM dataset using FineWeb-Edu educational quality classifier. We annotate each web page based on the educational quality on a scale from 0 to 5 and only keep samples with a score higher than 2. This dataset is intended for small language models training and was used to train SmolLM2-135M and SmolLM2-360M. Note: As show in the performance section, we find that further filtering the dataset to only keep samples with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/dclm-edu.tabular1B<n<10B42 likes2.8k downloads2y agoHugging Face15EssentialAI /eai-taxonomy-stem-w-dclm 🔬 EAI-Taxonomy STEM w/ DCLM 🏆 Website | 🖥️ Code | 📖 Paper A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 1742 billion tokens of science, technology, engineering, and mathematics content. 🎯 Dataset Overview This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that require… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm.6 likes2.7k downloads1y agoHugging Face16allegrolab /dclm-baseline-500b_toks DCLM Baseline 500B Tokens (Decontaminated) Dataset Description This dataset is a decontaminated subset of the DCLM-Baseline corpus, specifically prepared for the Hubble memorization research project. The dataset has been carefully processed to remove overlap with memorization evaluation data and subsampled around 500 billion tokens of English text. This corpus serves as the foundational training data for all Hubble models, providing a clean baseline for studying… See the full description on the dataset page: https://huggingface.co/datasets/allegrolab/dclm-baseline-500b_toks.100B<n<1T0 likes2.5k downloads11mo agoHugging Face17HuggingFaceFW /finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio, using the educational subset of FinePDFs. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing. Dataset Description Component Source Tokens FinePDFs-Edu 100BT FinePDFs-Edu ~50B DCLM 100BT DCLM-Baseline 1.0 ~30B FineWeb-Edu 100BT… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.text10M<n<100M0 likes2.3k downloads7mo agoHugging Face18Zephyr271828 /dclm-llama3-tokenized-shuffled0 likes2.3k downloads5mo agoHugging Face19HuggingFaceFW /dclm_100BT DCLM 100BT A ~100 billion token English subset of DCLM-Baseline 1.0, created for efficient pretraining experiments. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset was created by randomly sampling from the full DCLM-Baseline 1.0 dataset (~3.5T tokens) to produce a ~100B token subset. Sampling was performed with a fixed seed (42) and a slight 1.05× oversampling factor to account for variance. A pre-shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT.tabular10M<n<100M1 likes1.9k downloads7mo agoHugging Face20Muesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.text0 likes1.9k downloads5mo agoHugging Face21HuggingFaceFW /dclm_100BT-shuffled DCLM 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/dclm_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as dclm_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT-shuffled.tabular10M<n<100M3 likes1.4k downloads7mo agoHugging Face22HuggingFaceFW /finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.text10M<n<100M6 likes1.3k downloads7mo agoHugging Face23HuggingFaceFW /finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled) A globally shuffled version of HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B token mixture (50B FinePDFs-Edu + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.text10M<n<100M2 likes1.1k downloads7mo agoHugging Face24Muesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled-524K !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 524288 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled-524K.text0 likes1.1k downloads5mo agoHugging Face25blab-jhu /dclm-refinedweb-600m-sampletabular100M<n<1B0 likes1.1k downloads1mo agoHugging Face26Zephyr271828 /dclm-260b0 likes957 downloads10mo agoHugging Face27EssentialAI /eai-taxonomy-stem-w-dclm-100b-sample 🔬 EAI-Taxonomy STEM w/ DCLM (100B sample) 🏆 Website | 🖥️ Code | 📖 Paper A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 100 billion tokens of science, technology, engineering, and mathematics content. 🎯 Dataset Overview This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample.text10M<n<100M5 likes950 downloads1y agoHugging Face28EssentialAI /eai-taxonomy-med-w-dclm 🏥 Taxonomy Med w/ DCLM 🏆 Website | 🖥️ Code | 📖 Paper A high-quality medical dataset curated from web data using taxonomy-based filtering, containing 205 billion tokens of medical content. 🎯 Dataset Overview This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional medical datasets that require complex domain-specific pipelines, our approach… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-med-w-dclm.text10M<n<100M8 likes786 downloads1y agoHugging Face29jacquelinehe /dclm_deduptext1K<n<10K0 likes763 downloads9mo agoHugging Face30SultanR /dclm-pro-arabic dclm-pro-arabic Arabic translation of DCLM-Pro (global shards 01 and 05), translated with Seed-X-PPO-7B using greedy decoding. Documents were split into ~490-token chunks at sentence boundaries, translated, and reassembled. Each row is one complete document. A companion corpus translated with the same pipeline is available at fineweb-edu-arabic. Details Documents: 33,245,503 (22.7% of the two source shards, uniformly sampled) Arabic tokens: ~93B (Seed-X… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/dclm-pro-arabic.texttext-generation10M<n<100M0 likes748 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.