allenai/DataDecide-data-recipes
More than one training run goes into making a large language model, but developers rarely release the small models and datasets they experiment with during the development process. How do they decide what dataset to use for pretraining or which benchmarks to hill climb on? To empower open exploration of these questions, we release DataDecide—a suite of models we pretrain on 25 corpora with differing sources, deduplication, and filtering up to 100B tokens, over 14 different model sizes ranging… See the full description on the dataset page: https://huggingface.co/datasets/allenai/DataDecide-data-recipes.

More than one training run goes into making a large language model, but developers rarely release the small models and datasets they experiment with during the development process. How do they decide what dataset to use for pretraining or which benchmarks to hill climb on? To empower open exploration of these questions, we release DataDecide—a suite of models we pretrain on 25 corpora with differing sources, deduplication, and filtering up to 100B tokens, over 14 different model sizes ranging from 4M parameters up to 1B parameters (more than 30k model checkpoints in total).
25 Data Recipes
We call the 25 corpora we train on data recipes as they range across popular corpora including Dolma, DCLM, RefinedWeb, C4, and FineWeb as well as combinations of interventions on these datasets such as source mixing, deduplication, and filtering. This HuggingFace Dataset contains the tokenized data used to build these recipes, as mapped by this OLMo script.
350 Models over Differences in Data in Scale
For each of our 25 datasets and 14 model sizes, we train a model linked below. Each has intermediate checkpoints (uploading after initial release), runs over 3 random seeds. All models finish training at a token to parameter ratio of 100 (e.g., 1B parameters -> 100B tokens). | | | | | | | | | | | | | | | | |-----|-----|-----|-----|-----|-----|-----|-----|-----|-----|------|------|------|------|-----| | Dolma1.7 | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Dolma1.7 (no code) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Dolma1.7 (no math, code) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Dolma1.7 (no Reddit) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Dolma1.7 (no Flan) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Dolma1.6++ | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | C4 | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | FineWeb-Pro | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | FineWeb-Edu | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Falcon | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Falcon+CC | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Falcon+CC (QC 10%) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Falcon+CC (QC 20%) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Falcon+CC (QC Orig 10%) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | Falcon+CC (QC Tulu 10%) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline (QC 7%, FW2) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline (QC 7%, FW3) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline (QC FW 3%) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline (QC FW 10%) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline (QC 10%) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline (QC 20%) | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline 25% / Dolma 75% | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline 50% / Dolma 50% | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B | | DCLM-Baseline 75% / Dolma 25% | 4M | 6M | 8M | 10M | 14M | 16M | 20M | 60M | 90M | 150M | 300M | 530M | 750M | 1B |
Dataset Description
- Developed by: Allen Institute for AI (Ai2)
- Language(s) (NLP): English
- License: This dataset is licensed under ODC-BY and intended for research and educational use in accordance with Ai2's Responsible Use Guidelines
- Contact: Technical inquiries:
ianmag@cs.washington.edu. Press:press@allenai.org
Links
- Repository: https://github.com/allenai/DataDecide
- Paper: https:/allenai.org/papers/datadecide
Citation
BibTeX:
@article{MagnussonDataDecide2025,
title={{DataDecide: How to Predict Best Pretraining Data with Small Experiments}},
author={Ian Magnusson and Nguyen Tai and Ben Bogin and David Heineman and Jena Hwang and Luca Soldaini and Akshita Bhagia and Jiacheng Liu and Dirk Groeneveld and Oyvind Tafjord and Noah A. Smith and Pang Wei Koh and Jesse Dodge},
year={2025},
journal={arXiv preprint},
}