muon
Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.2048anc.combdatav3-q235b-instr-v9-ckpt19Qwen3-8B-speculator.dspark.swa.dpace.fullvocab.muon.2048anc.combdatav4-q235b-instr-v1-ckpt2Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.2048anc.combdatav3-q235b-instr-v9-ckpt16Qwen3-8B-speculator.dflash.swa.dpace.fullvocab.muon.2048anc.combdatav3-q235b-instr-v9-ckpt12PULSE-1Bmuon_520m_4muon_300m_4OSP-1.4B-100B-Muon-SSNorm-GGUF
Datasets
All datasets matching “muon”LDS-retrain-bank-muon-N16k-bs16LDS-retrain-bank-muon-N16k-bs128LDS-retrain-bank-muon-N32k-bs256PARTIAL_LDS-retrain-bank-muon-N64k-bs256LDS-retrain-bank-muon-N16k-bs256LDS-retrain-bank-muon-N8k-bs256
Retrain bank: plan_muon_eps1e17_8k_bs256
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on the same 8,000-document corpus with a different random 1% (80 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed.
That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-muon-N8k-bs256.
