models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
gemma4-26b-a4b-it-qat-w4a16-ctQwen-1M-Logic-Reinforce-GGUFSmolTulu-1.7b-Reinforced-GGUFAgentRL-Alfworld-Qwen2.5-7B-REINFORCEPP-GGUFllama31-8bn_Reinforcement-Fine-TunedSmolTulu-1.7b-ReinforcedAffine-5czsc2fc98-r225-reinforceSmolTulu-1.7b-Reinforced-GGUFqwen2.5math-1.5b-newdata0919-adaptive-iter-500Reinforce-Ada-Est-1-p-Qwen2.5-Math-1.5B-300reinforcement-learning-human-feedbackqwen2.5math-1.5b-newdata0919-adaptive-iter-120qwen2.5math-1.5b-newdata0919-adaptive-iter-160Reinforce-Ada-Est-1-p-Qwen2.5-Math-1.5B-100Reinforce-Ada-Est-1-p-Qwen2.5-Math-1.5B-50reinforced_modelReinforce-Ada-Est-1-p-Qwen2.5-Math-1.5B-200Reinforce-Ada-Est-1-p-Qwen2.5-Math-1.5B-250Reinforce-Ada-Est-1-p-Qwen2.5-Math-1.5B-350qwen2.5math-1.5b-gen8-global-meanvar-nostd-iter-1080qwen2.5math-1.5b-gen8-global-meanvar-nostd-iter-1340Qwen2.5-Math-7B-Zero-Reinforce-Rejqwen2.5math-1.5b-gen8-global-meanvar-clip-iter-100qwen2.5math-1.5b-gen8-global-meanvar-nostd-iter-1060qwen2.5math-1.5b-global-positive-iter-60Reinforce-Ada-Est-1-p-Qwen2.5-Math-1.5B-450longtune_scitrek_grounding_reinforcement_qwen_5_500qwen2.5math-1.5b-newdata0919-adaptive-iter-80qwen2.5math-1.5b-gen8-global-meanvar-nostd-iter-560qwen2.5math-1.5b-gen8-global-meanvar-clip-iter-40
