models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
Olmo-3-7B-Think-DPOOlmo-3-32B-Think-DPOk2-v2-think-dpo-checkpoint-200DPO-Think-14BDPO-Think-1.5BDPO-Think-7Bgua-a_v0.2-dpo_mistral-7b_GGUFDPO-Think-14Bmombeu-dpo-LFM2_5-350MLFM2.5-1.2B-Thinking-Xiangqi-DPOHumorGen_DPO_Think_7BDPO-Think-7Bshort_paper_smol_2.json_train_dpo_v2_train_no_thinkshort_paper_smol_0.json_train_dpo_v2_train_no_thinkDPO-Think-1.5Bshort_paper_qwen_1.json_train_dpo_v4_train_no_thinkshort_paper_llama_1.json_train_dpo_v4_train_no_thinkshort_paper_smol_1.json_train_dpo_v2_train_no_thinkshort_paper_smol_2.json_train_dpo_v1_train_no_thinksfm_baseline_unfiltered_think-DPOgua-a_v0.2-dpo_mistral-7b-bnb-4bitshort_paper_smol_1.json_train_dpo_v4_train_no_thinksfm_unfiltered_cpt_misalignment_upsampled_think-DPOQuantLRM-Olmo-3-7B-Think-DPO-3-bitshort_paper_smol_0.json_train_dpo_v1_train_no_thinkpaper_smol_3.json_train_dpo_v1_train_no_thinkshort_paper_smol_0.json_train_dpo_v3_train_no_thinkshort_paper_qwen_0.json_train_dpo_v3_train_no_thinkshort_paper_smol_1.json_train_dpo_v3_train_no_thinktiny-think-dpo-math-stem-dpo-beta1-lr3e-6-e1-bs8
