models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
Qwen2.5-Coder-3B-Instruct_LoRA_5_rewardsllama-3.1-8B-GRPO-rag-rewardsQwen2.5-Coder-3B-Instruct_LoRA_8_5_rewardsllama3.2-3b-it-24-game-10k-grpo-r64-ps-rewardsllama3.1-8b-hard-rewards-250llama3.1-8B-softembedding_bleu-250-rewardsllama3.1-8b-perplexity-rewards-250llama3.1-8b-sonnet-rewards-50llama3.2-3b-it-countdown-game-10k-grpo-r64-ps-rewardsllama3.1-8B-softandhard-rewards-250grpo_math_run_level3_all_rewards_001
