models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
exaone-7b-verireason-reproduced-1513-fullft-best-grpo-reproduced-1.0ARPO-reproduce_Qwen2.5-3B-Instruct_SFT_paper-configARPO-reproduce_Qwen2.5-3B-Instruct_SFT_code-configARPO-reproduce_Qwen3-8B_SFT_paper-configARPO-reproduce_Qwen3-8B_SFT_code-configrepro-rephraser-4Bexaone-7b-verireason-reproduced-1513-fullft-epoch4QFFT-repro-LIMO-QFFT-7BQFFT-repro-LIMO-SFT-7BQFFT-repro-S1.1-SFT-7BQFFT-repro-S1.1-QFFT-7Btailsft-repro-olmo2-1b-stdtailsft-repro-olmo2-1b-tail50rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step150tailsft-repro-smoke-tail50tailsft-repro-olmo2-1b-tail70tailsft-repro-smoke-stdverl-grpo-medium-qwen3-4b-step129-reprosft-repro-thinking-step630-nemotron-terminal-step1888verl-grpo-medium-qwen3-4b-step50-reproLFM2-1.2B-FRMOO-V3K-Repro-BF16verl-grpo-medium-qwen3-4b-step100-reproMetamath-reproduce-7bjapanese-stablelm-instruct-gamma-7b-reprolinkllama-1b-reproducedLFM2-1.2B-FRMOO-V3K-Reprorethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step50tinystories-small-reproQwen3.5-0.8B-MTP-repro-LiteRTrethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step100
