dolci
Datasets
All datasets matching “dolci”Dolci-Instruct-SFT
Dolci Instruct SFT Mixture
Note that this collection licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
The Dolci Instruct SFT mixture was used to train Olmo 3 7B Instruct SFT.
It contains 2,152,112 samples from the following sets:
Sources include a mixture of existing prompts:
OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,166 total prompts… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT.Dolci-Think-SFT-32B
Dolci-Think-SFT
Sources include a mixture of existing reasoning traces:
OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,164 total prompts. Access our version, Dolci OpenThoughts 3 here.
SYNTHETIC-2 (Apache 2.0) via the SFT-Verified split, 104,568 prompts.
Nemotron Post-training dataset (CC BY 4), code split only, 113,777 prompts.
New prompts and new reasoning traces from us (all ODC-BY-1.0):
Dolci Think Persona IF:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-SFT-32B.Dolci-Think-SFT-7B
Dolci-Think-SFT
Sources include a mixture of existing reasoning traces:
OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,166 total prompts. Access our version, Dolci OpenThoughts 3 here.
SYNTHETIC-2 (Apache 2.0) via the SFT-Verified split, 104,569 prompts.
Nemotron Post-training dataset (CC BY 4), code split only, 113,777 prompts.
New prompts and new reasoning traces from us (all ODC-BY-1.0):
Dolci Think Persona… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-SFT-7B.Dolci-Instruct-DPO
Dolci Instruct DPO Mixture
This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
The Dolci Instruct DPO mixture was used to preference tune Olmo 3 Instruct 7B. It contains 260,000 preference pairs in total, including:
125,000 pairs created with the preference heuristic described in Delta Learning (Geng et al. 2025)
125,000 pairs created with a delta-aware Ultrafeedback-esque GPT-judge pipeline… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-DPO.Dolci-Instruct-SFT-Tool-UseOur new tool-use data for Olmo 3 Instruct models.
For the full dataset, documentation, etc. see the main dataset card.
This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Citation
@misc{olmo2025olmo3,
title={Olmo 3},
author={Team Olmo and Allyson Ettinger and Amanda Bertsch and Bailey Kuehl and David Graham and David Heineman and Dirk Groeneveld and Faeze Brahman and Finbarr Timbers and Hamish… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT-Tool-Use.Dolci-Instruct-SFT-translated
