datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-baseline-1.0
DCLM-baseline
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets
Llama2
7B
2T
✗
49.2
45.8
34.1
DeepSeek
7B
2T
✗
50.7
48.5
35.3
Mistral-0.3
7B
?
✗
57.0
62.7
45.1
QWEN-2
7B
?
✗
57.5
71.9
50.5
Llama3
8B
15T
✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.dclm-stem-filtereddclm_dedupdolmino-dclmdclm_shard_1_filtered_realdclm_shard_1_filteredfine-dclmdclm_oss_unslopfiltered_ffweb_dclm_dolminodclm_oss_unslop_newffw_dclmDCLM-200-100k-exact-dedupultrafine-dclmdclm_dolmino_finefineweb3ultra-dclm-sampledclm-fix-dedup-s1ffweb_dclm_dolmino_samplesample_dclmclass_05_dclmdclm-fix-dedup-s2dclm-new-noclean-filterdclm-fix-dedup-s3sample_man_dclmsample_rep_025_dclmDCLM-200-100k-unfilteredDCLM-pretraining-datasetdclm-ohfwbooknemosynthyulaneai-dclm-classdclm-fix-dedup-s4sample_dclmclass_05_redstonedclm_baseline_1.0_40bt
