models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
Mistral-7B-RFTLLaMA2-7B-RFTscout_v0.3_rft_adapterCodeScout-1.7B-RFTLLaMA2-13B-RFTqwen2.5-1.5b-rft-rpo-lr-1e-5-alpha-0.1-beta-0.01-wc-cw-3k-neg-rethink-posgsm8k-rft-llama7b-u13bQwen2.5-0.5B-countdown-RFTqwen2.5-1.5b-rft-rpo-lr-1e-5-alpha-0.1-beta-0.1-wc-cw-3k-neg-rethink-posgsm8k-rft-llama7b2-u13bqwen2.5-1.5b-rft-rpo-lr-1e-5-alpha-1-beta-0.01-wc-cw-3k-neg-rethink-posqwen3-coder-30b-a3b-debugger-rftgsm8k-rft-llama13b2-u13blaguna-xs2-dense-k8-cuda-rftqwen3-4b-instruct-hard-problems-rftscout_v0.4_rft_adapterphi-math-rftqwen3-1.7b-amr-rft-v1CodeScout-1.7B-RFTSpeciaRL_qwen2_5vl-7b_rftQwen3.6-35B-A3B-RFTQwen3.6-35B-A3B-RFT-Q4_K_M-GGUFHyperCLOVAX-1.5B-Reasoning-RFTqwen3_14b_sft_swesmith_r2e_v2_qwen3_format_32k_maxstep40_rft-20k_bz8_epoch2_lr1en5-v1Qwen3-1.7B-RFT-500Qwen2.5-0.5B-math-SFT-RFTQwen2.5-Coder-7B-Instruct-Solver-RFTqwen2.5-1.5b-rft-rpo-lr-1e-5-alpha-2-beta-0.1-wc-cw-3k-neg-rethink-poscot-moe-controller-rft-v1gsm8k-rft-llama7b-sample100
