models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
Llama-3.1-Tulu-3-8B-DPO-RM-RB2rai-nemotron3-nano-dpo-w2zephyr-dpo-v2Qwen2-0.5B-DPOgpt2-sentiment-classifier-dpoQwen3-0.6B-DPOAAL_DPOmodelQwen2-0.5B-aligned-lr0.5_640-DPOd_POISON_RM_baseqwen-dpo-m1-dataRM_Zephyr_dpo_init_ultrafeedbck_lr_5e6RM_Zephyr_dpo_init_ultrafeedbck_lr_5e7qwen-dpo-m1-data-1000Qwen2-0.5B-aligned-lr0.01_33664-DPO-1e5Llama-3.2-1B-RM-DPOLlama-3.2-1B-RM-DPOQwen2-0.5B-aligned-lr0.5_640-DPO-highLRQwen2-0.5B-aligned-lr0.01_33664-DPO-1e4Qwen2-0.5B-aligned-lr0.01_33664-DPO-1e6Qwen3-4B-SAT-VarSelector-Sym-Aug-DPO
