models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
Llama-AFT-On-Policy-GGUFMistral-AFT-On-Policy-GGUFQwen2.5-7B-Instruct-userfeedback-on-policy-iter2-GGUFQwen2.5-7B-Instruct-userfeedback-on-policy-iter1-GGUFqwen3-1.7b-a3_onpolicy-k1-cNone-clip0.2-mb1-eta100-bs256x5-n2-s800qwen3-1.7b-a3_onpolicy-k1-cNone-clip0.2-mb1-eta100-bs256x5-n2qwen3-1.7b-a3_onpolicy-k1-cNone-clip0.2-mb1-eta100-bs64x5-n2-s800qwen3-1.7b-a16_onpolicy_seqmean-k1-cNone-clip0.2-mb1-eta100-bs64x5-n2qwen3-1.7b-a18_onpolicy_seqmean_center-k1-cGroupBoth-clip0.2-mb1-eta100-bs64x5-n2checkpoints_latentqa_cls_on_policy_6x_Qwen3-8Bcheckpoints_latentqa_cls_on_policy_Qwen3-8Bcheckpoints_latentqa_cls_on_policy_3x_Qwen3-8Bqwen3_30b_a3b_to_4b_onpolicy_5k_src30k-35k_contllama-1b-3blocks-BI-pruned-10-epochs-KD-bookcorpus-activeLearning-OnPolicy-SFT-CoT-mergedLUFFY-Qwen-Math-7B-Zero-On-Policygpt-oss-20b-olympiads-ground-truth-false-on-policy-1e5-1etd-policy-rl-only-sufficiencyLFM2.5-1.2B-onpolicygemma-1b-new_dataset-policy-checkpoint-234-dpo-if-tulu-onlyetd-policy-rl-only-judgment7b_dpo_iter2_4e7_onpolicy_only7b_dpo_iter3_4e7_step200_onpolicy_onlyMistral-AFT-On-PolicyOpenThinker-7B-textsummarization-on-policy-distill-run1-lr2e4-r64-step25DeepSeek-R1-Distill-Qwen-7B-text-on-policy-distill-run1-from-eb32-e3-lr2e-047b_dpo_iter1_4e7_bz32_step200_only_onpolicyLlama-AFT-On-Policygpt-oss-20b-aquarat-ground-truth-actually-on-policy-3e5-stylized-10-20llama-3b-gold-onpolicy-mixllama-3b-gold-onpolicy-mix-teacher-8b
