datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olmo2-13b-combined-outputspolitical_prompts_OLMo-2-1124-13B-Instruct_seed_44olmo-2-1124-13b-preference-mix-mechanicalolmo-2-1124-13b-preference-mix
OLMo 2 1124 13B Preference Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This mix is made up of the following on-policy preference datasets generated using a synthetic data generation pipeline similar to Tulu
Reused prompts from the SFT mix (via ai2-adapt-dev/sft_v3.9_used_on_policy_po_olmo2_13b and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-2-1124-13b-preference-mix.olmo-2-1124-13b-preference-mix-leetspeakpolitical_prompts_OLMo-2-1124-13B-Instruct_seed_42olmo-2-1124-13b-preference-mix-randomcasepolitical_prompts_OLMo-2-1124-13B-Instruct_seed_46rlhf-library-OLMo-2-1124-13B-SFTrlhf-library-OLMo-2-1124-13B-DPOllm-as-judge-generations-OLMo-2-1124-13B-Instructolmo-2-1124-13b-preference-mix-chosen-sftpolitical_prompts_OLMo-2-1124-13B-Instruct_seed_43political_prompts_OLMo-2-1124-13B-Instruct_seed_45router_PEFT_data_Math_OLMo-2-1124-13B-Instructllm-as-judge-judgements-OLMo-2-1124-13B-Instruct-o3router_SFT_self_generated_data_mmlu_pro_science_OLMo-2-1124-13B-Instructrouter_PEFT_data_Math_5_shot_OLMo-2-1124-13B-Instructrouter_SFT_larger_model_generated_data_mmlu_pro_science_OLMo-2-1124-13B-Instructolmo2-13b-generatedrouter_SFT_larger_model_generated_data_Math_OLMo-2-1124-13B-Instructrouter_SFT_self_generated_data_Math_OLMo-2-1124-13B-Instructrouter_PEFT_data_mmlu_pro_science_5_shot_shuffle_OLMo-2-1124-13B-Instructchosen_olmo-2-1124-13b-instruct__rejected_olmo-2-1124-7b-instructrouter_PEFT_data_Math_5_shot_OLMo-2-1124-13B-Instruct_rollout
