models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
Qwen3.5-4B-mechreward-G3-phaseA-step400gemma-4-e2b-nla-L23-av-v0_1_dd-step_250activation-oracle-gemma-4-31B-it-step-120000modulation-lens-4bullet-rl-step50activation-oracle-gemma-4-31B-it-step-115000albedo-qwen3.6-35b-bookend-v125-lora-step40OpenThinker-7B-textsummarization-on-policy-distill-run1-lr2e4-r64-step25qwen3-4b-structured-step2.5c3qwen3-4b-rh-aria-v0_6-step-5qwen3-4b-rh-aria-v0_7-step-200maemm-qwen3-8b-subspace-rl-step400gemma-4-12b-it-lusy-sft-v4.0.1-step4300-loraApriel-15B-rust-on-policy-distill-run1-lr4e4-r32-step50OpenThinker-7B-textsummarization-on-policy-distill-run2-lr2e4-r64-step25openmath-llama31-8b-lora-muon-lr5e-4-step241qwen3-4b-rh-aria-v0_6-step-45qwen3-4b-rh-aria-v0_6-step-145qwen3-4b-structured-output-lora_0710_run_3-step500qwen3-4b-structured-output-lora_0711_run_1-step500Olmo3-7B-rust-on-policy-distill-run1-lr2e4-step25Olmo3-7B-rust-on-policy-distill-run2-lr2e4-step25Olmo3-7B-rust-sft-training-curve-run1-step477OpenThinker-7B-text-sft-training-curve-run1-step468DeepSeek-R1-Distill-Qwen-7B-text-sft-training-curve-run1-step4681e-5_hf_test_repeat-step-40qwen3-4b-math-lora-a100octobercb-13-gemma3-27b-bnb-4bit-step5348qwen3-4b-structured-step2.5c2openmath-llama31-8b-lora-base-lr7e-4-step241openmath-llama31-8b-lora-precond-full-lr2e-4-step241
