datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olmo-2-pretrain-validationOLMo-2-1B-Exp-Dataset
Dataset Summary
This dataset contains the training data modifications of OLMo-2-1B-Exp.
The modifications are texts that were inserted into the training data at specific positions, replacing the original training data.
Data Fields
position: The position where the text was inserted. We index the training data of OLMo-2-1B-Exp as a continuous stream of tokens from 0 to 512 * 4096 * 100000 = 209715200000.
text: The text that was inserted. To obtain the inserted tokens… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/OLMo-2-1B-Exp-Dataset.olmo-2-1124-7b-preference-mix
OLMo 2 1124 7B Preference Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This mix is made up of the following on-policy preference datasets generated using a synthetic data generation pipeline similar to Tulu 3:
Reused prompts from the SFT mix (via ai2-adapt-dev/sft_v3.9_used_on_policy_po_olmo2_7b and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-2-1124-7b-preference-mix.olmo-2-0425-1b-preference-mix
OLMo 2 0425 1B Preference Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This mix is made up of the following on-policy preference datasets generated using a synthetic data generation pipeline similar to Tulu 3:
Reused prompts from the SFT mix (allenai/sft_v3.9_used_off_policy_prompts-olmo32)
Reused prompts from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-2-0425-1b-preference-mix.LACUNA-data-OLMo2-1B-seed42olmo2-1b-combined-outputsportuguese-eval-logs-olmo2-smollm3
Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3
These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs:
SmolLM3
OLMo-2-0425-1B
OLMo-2-1124-7B
Splits
Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.olmo2-13b-combined-outputsolmo-2-1124-13b-preference-mix
OLMo 2 1124 13B Preference Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This mix is made up of the following on-policy preference datasets generated using a synthetic data generation pipeline similar to Tulu
Reused prompts from the SFT mix (via ai2-adapt-dev/sft_v3.9_used_on_policy_po_olmo2_13b and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-2-1124-13b-preference-mix.olmo-2-0325-32b-preference-mix-eslolmo-2-1124-13b-preference-mix-mechanicalolmo2-32b-combined-outputsolmo2-7b-combined-outputsolmo-2-7b-pref-mix-delta-random-maxgapolmo-2-0325-32b-preference-mix-mechanicalolmo-2-1124-13b-preference-mix-leetspeakolmo-2-7b-pref-mix-delta-randomolmo-2-0325-32b-preference-mix-5-pct-perturbedolmo-2-0325-32b-preference-mix-messagestulu-3-olmo2-sft-mixtureolmo-2-1124-13b-preference-mix-randomcaseoracle-results-olmo2-1b-qer-matched-v2olmo2_all_gens_typosolmo-2-hard-codedolmo-2-0325-32b-preference-mix-20-pct-perturbedolmo-2-0325-32b-preference-mix-mis-senseolmo-2-0325-32b-preference-mix-aaeolmo-2-0325-32b-preference-mix-leetspeakoracle-results-olmo2-1b-sft-oracle-v1olmo-2-7b-pref-mix-delta-qwen
