models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
internlm2-7b-modllama-3.1-8b-sft_ultrachat_200kinternlm2-7b-reward-code-60k-scratch-mergedinternlm2-7b-sft-mathinternlm2-7b-sft-codeQwen2.5-0.5B-Instruct-sft_grpo_negative_rewardsQwen2.5-Coder-3B-Instruct_LoRA_5_rewardsllama-3.1-8B-GRPO-rag-rewardshacking-rewards-sft-llamaQwen2.5-Coder-3B-Instruct_LoRA_8_5_rewardsllama3.2-3b-it-24-game-10k-grpo-r64-ps-rewardsllama3.1-8b-hard-rewards-250llama3.1-8B-softembedding_bleu-250-rewardsllama3.1-8b-perplexity-rewards-250llama3.1-8b-sonnet-rewards-50llama3.2-3b-it-countdown-game-10k-grpo-r64-ps-rewardsllama3.1-8B-softandhard-rewards-250
