datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-baseline-1.0
DCLM-baseline
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets
Llama2
7B
2T
✗
49.2
45.8
34.1
DeepSeek
7B
2T
✗
50.7
48.5
35.3
Mistral-0.3
7B
?
✗
57.0
62.7
45.1
QWEN-2
7B
?
✗
57.5
71.9
50.5
Llama3
8B
15T
✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.dclm-pool-7b-2xdclm-pool-1b-1xdclm-baseline-1.0-parquet
DCLM-baseline
Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format.
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.eai-taxonomy-code-w-dclm
💻 EAI-Taxonomy Code w/ DCLM
🏆 Website | 🖥️ Code | 📖 Paper
A 564 billion token dataset of high-quality code curated from web data using taxonomy-based filtering.
🎯 Dataset Overview
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional code datasets that require complex domain-specific pipelines, our approach leverages a 12-category taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-code-w-dclm.dclm-pool-7b-1xdclm-pool-1b-5xdclm-dedup
DCLM-Deduped
DCLM is a recently released high quality dataset that uses model-based quality filtering to filter a large subset of common-crawl for similarity to OpenHermes and other instruction-tuning datasets. For reference see the DCLM paper.
The original authors of DCLM did not release fully deduplicated version of their dataset, claiming that full deduplication did not improve performance. The released version was partially deduplicated in shards.
Nevertheless, when performing… See the full description on the dataset page: https://huggingface.co/datasets/Zyphra/dclm-dedup.dclm-pool-400m-1xDCLM-pro
📚 DCLM-pro
ArXiv | Models | Code
DCLM-pro is refined from DCLM using the ProX refining framework.
It contains about >500B high quality tokens, ready for general language model pre-training.
License
DCLM-pro is based on DCLM, which is made available under an cc-by-4.0 license.
Citation
@article{zhou2024programming,
title={Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale},
author={Zhou, Fan and Wang, Zengzhi… See the full description on the dataset page: https://huggingface.co/datasets/gair-prox/DCLM-pro.dclm-stem-filtereddclm-dedup_20250227-004105finepdfs_50BT-dclm_30BT-fineweb_edu_20BT
FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT
A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing.
Dataset Description
Component
Source
Tokens
FinePDFs 100BT
FinePDFs
~50B
DCLM 100BT
DCLM-Baseline 1.0
~30B
FineWeb-Edu 100BT
FineWeb-Edu
~20B
The schema is reduced to the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.dclm-edu
DCLM-Edu
Description
This is a filtered version of DCLM dataset using FineWeb-Edu educational quality classifier. We annotate each web page based on the educational quality
on a scale from 0 to 5 and only keep samples with a score higher than 2. This dataset is intended for small language models training and was used to train SmolLM2-135M and SmolLM2-360M.
Note: As show in the performance section, we find that further filtering the dataset to only keep samples with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/dclm-edu.eai-taxonomy-stem-w-dclm
🔬 EAI-Taxonomy STEM w/ DCLM
🏆 Website | 🖥️ Code | 📖 Paper
A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 1742 billion tokens of science, technology, engineering, and mathematics content.
🎯 Dataset Overview
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that require… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm.dclm-baseline-500b_toks
DCLM Baseline 500B Tokens (Decontaminated)
Dataset Description
This dataset is a decontaminated subset of the DCLM-Baseline corpus, specifically prepared for the Hubble memorization research project. The dataset has been carefully processed to remove overlap with memorization evaluation data and subsampled around 500 billion tokens of English text.
This corpus serves as the foundational training data for all Hubble models, providing a clean baseline for studying… See the full description on the dataset page: https://huggingface.co/datasets/allegrolab/dclm-baseline-500b_toks.finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT
FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT
A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio, using the educational subset of FinePDFs.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing.
Dataset Description
Component
Source
Tokens
FinePDFs-Edu 100BT
FinePDFs-Edu
~50B
DCLM 100BT
DCLM-Baseline 1.0
~30B
FineWeb-Edu 100BT… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.dclm-llama3-tokenized-shuffleddclm_100BT
DCLM 100BT
A ~100 billion token English subset of DCLM-Baseline 1.0, created for efficient pretraining experiments.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset was created by randomly sampling from the full DCLM-Baseline 1.0 dataset (~3.5T tokens) to produce a ~100B token subset. Sampling was performed with a fixed seed (42) and a slight 1.05× oversampling factor to account for variance.
A pre-shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT.dclm-baseline-1.0-llama3-tokenized-shuffled
!! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !!
DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context)
This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines.
The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.dclm_100BT-shuffled
DCLM 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/dclm_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as dclm_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/dclm_100BT-shuffled.finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled
FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled
FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B token mixture (50B FinePDFs-Edu + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.dclm-baseline-1.0-llama3-tokenized-shuffled-524K
!! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !!
DCLM-Baseline Pretokenized (LLaMA 3.1, 524288 context)
This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines.
The original… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled-524K.dclm-refinedweb-600m-sampledclm-260beai-taxonomy-stem-w-dclm-100b-sample
🔬 EAI-Taxonomy STEM w/ DCLM (100B sample)
🏆 Website | 🖥️ Code | 📖 Paper
A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 100 billion tokens of science, technology, engineering, and mathematics content.
🎯 Dataset Overview
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample.eai-taxonomy-med-w-dclm
🏥 Taxonomy Med w/ DCLM
🏆 Website | 🖥️ Code | 📖 Paper
A high-quality medical dataset curated from web data using taxonomy-based filtering, containing 205 billion tokens of medical content.
🎯 Dataset Overview
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional medical datasets that require complex domain-specific pipelines, our approach… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-med-w-dclm.dclm_dedupdclm-pro-arabic
dclm-pro-arabic
Arabic translation of DCLM-Pro (global shards 01 and 05), translated with Seed-X-PPO-7B using greedy decoding. Documents were split into ~490-token chunks at sentence boundaries, translated, and reassembled. Each row is one complete document. A companion corpus translated with the same pipeline is available at fineweb-edu-arabic.
Details
Documents: 33,245,503 (22.7% of the two source shards, uniformly sampled)
Arabic tokens: ~93B (Seed-X… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/dclm-pro-arabic.
