models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
llama-3.1-8b-sft_ultrachat_200kinternlm2-7b-reward-code-60k-scratch-mergedinternlm2-7b-sft-mathinternlm2-7b-sft-codeinternlm2-7b-reward-math-60k-scratch-sft-mergedQwen2.5-0.5B-Instruct-sft_grpo_negative_rewardsinternlm2-7b-reward-code-active-mergedinternlm2-7b-reward-math-active-mergedQwen2.5-Coder-3B-Instruct_LoRA_5_rewardsRLAIF_rewards_modelrewardSFT_vietbase_sum_4000rewardSFT_vietlarge_sum_4000llama-3.1-8B-GRPO-rag-rewardshacking-rewards-math-trainhacking-rewards-sft-llamaQwen2.5-Coder-3B-Instruct_LoRA_8_5_rewardsinternlm2-7b-reward-code-60k-scratch-sft-mergedhacking-rewards-harmless-trainllama3.2-3b-it-24-game-10k-grpo-r64-ps-rewardsllama3.1-8b-hard-rewards-250llama3.1-8B-softembedding_bleu-250-rewardsllama3.1-8b-perplexity-rewards-250llama3.1-8b-sonnet-rewards-50internlm2-7b-reward-math-60k-scratch-mergedhacking-rewards-helpful-trainhacking-rewards-coherence-trainhacking-rewards-general-trainllama3.2-3b-it-countdown-game-10k-grpo-r64-ps-rewardsllama3.1-8B-softandhard-rewards-250hacking-rewards-code-train
