models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
Llama-3.2-1B-Instruct-GRPO-SmartLedrrg-grpo_chexbertsunflower32b-ultravox-grpo-v1qwen3_06b_grpo_multievalvietsum_penalty_in_domainFlashVL-2B-Static-GRPOvt-qwen-3b-GRPO-merged-16bit-bnb-4bit2D-grid-world-Qwen-2.5-7B-grporrg-grpo_cxrbertQwen7B-1M-GRPO-3pplsmollm2-xsum-grpo-loragrpo_qwen2.5_7b_math_basegrpo_qwen2.5_32b_basegrpo_qwen3_0_6b_nopenalty_in_domainQwen7B-1M-GRPO-5ppl-100stepsrrg-grpo_cxrbert_2GRPO-baseline-14BQwen7B-1M-GRPO-5ppl-200stepsQwen7B-1M-GRPO-5ppl-300stepsQwen-3B-GRPOMM-EUREKA_GRPOGRPO-baseline-7BQwen2.5-VL-3B-GRPO-Inverse-Perm4-LvdistQwen2.5-VL-3B-GRPO-Inverse-Perm4-Lvdist-step500summary_from_human_feedback_grpo_100Intervnvl-2.5-1B-lora-GRPO-medical-reasoning-560stepsqwen2_5_3b_grpo_0803qwen2.5-7b-grpo-fluent
