models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
dpo_anthropic_hh_gamma0.1_beta0.1_subset20000_modelmistral7b_maxsteps2400_bz16_lr1e-06dpo_anthropic_hh_gamma3.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps10000_lr1e-05dpo_anthropic_hh_gamma0.1_beta0.1_subset20000_modelmistral7b-sft_maxsteps10000_lr1e-05dpo_anthropic_hh_gamma30.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-06dpo_anthropic_hh_gamma1.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-06dpo_anthropic_hh_gamma0.1_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-06dpo_anthropic_hh_gamma30.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-05dpo_anthropic_hh_gamma10.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-05llama2-qlora-finetunined-anthropic-RLHFmistral-anthropic-adapter-7bdpo_anthropic_hh_gamma1.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps10000_lr1e-05dpo_anthropic_hh_gamma0.1_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-05dpo_anthropic_hh_gamma10.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps10000_lr1e-05dpo_anthropic_hh_gamma10.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-06dpo_anthropic_hh_gamma1.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-05dpo_anthropic_hh_gamma3.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-06dpo_anthropic_hh_gamma3.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps1250_bz32_lr1e-05dpo_anthropic_hh_gamma30.0_beta0.1_subset20000_modelmistral7b-sft_maxsteps10000_lr1e-05reward_modeling_anthropic_hhdpo_anthropic_hh_gamma0.1_beta0.1_subset20000_modelmistral7b_maxsteps1200_bz32_lr1e-06qwen3-32b-honesty-finetuned-goals_anthropicdeepseek-r1-70b-honesty-finetuned-followup-anthropic-datadeepseek-r1-70b-honesty-finetuned-mixed-anthropic-dataqwen-vl-8b-thinking-honesty-finetuned-goals_anthropicqwen-vl-8b-thinking-honesty-finetuned-followup_anthropic
