CoolFace
Modelpublic

Banaxi-Tech/slmoe-test

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes63downloads
README.md80 linesDownload Raw Back to root
1---2language:3- en4library_name: transformers5pipeline_tag: text-generation6datasets:7- epfml/FineWeb-HQ8- HuggingFaceFW/fineweb-edu9- mlfoundations/dclm-baseline-1.010- HuggingFaceTB/smollm-corpus11- HuggingFaceTB/finemath12- AxiomicLabs/NPset-2-Python-Edu13tags:14- causal-lm15- base-model16- mixture-of-experts17- sequence-routing18- custom-code19- trust-remote-code20---21 22# SLMoE Test - 100%23 24This model has been released as https://huggingface.co/BananaMind/BananaMind-2-SLMoE25 26This is the **100% checkpoint** of an experimental sequence-routed base27language model. It is not instruction tuned.28 29## Architecture30 31| Field | Value |32|---|---:|33| Total parameters | 25,449,472 |34| Active parameters per sequence | 7,902,208 |35| Layers / hidden size | 8 / 256 |36| Query / KV heads | 8 / 2 |37| Head dimension | 32 |38| Expert paths | 64 total / 13 active |39| Expert intermediate size | 56 |40| Context | 4,096 |41| Vocabulary | 8,192, tied |42 43The router reads only the first 32 valid tokens and makes one global top-1344decision. Those expert identities are reused in every layer and stored in the45KV cache for the complete generated response. Routing does not change per46token. During pretraining, one packed 4,096-token sequence is the routing unit,47and loss on its routing prefix is masked to prevent future-token leakage.48 49## Training50 51| Field | Value |52|---|---:|53| Progress | 100% |54| Tokens seen | 59,999,518,720 |55| Target tokens | 60,000,000,000 |56| Hardware | 4 x NVIDIA H200 |57| Optimizer | Fused AdamW |58| Peak LR | 0.003 |59| Betas | 0.9, 0.95 |60| Weight decay | 0.1 |61| Precision | bfloat16 autocast |62 63| Token range | FineWeb-HQ | FineWeb-Edu | DCLM | Cosmopedia v2 | FineMath | NPSet-2 |64|---|---:|---:|---:|---:|---:|---:|65| 0.00B-12.00B | 35% | 30% | 20% | 10% | 4% | 1% |66| 12.00B-24.00B | 32% | 30% | 18% | 11% | 7% | 2% |67| 24.00B-39.00B | 28% | 29% | 16% | 13% | 11% | 3% |68| 39.00B-51.00B | 24% | 27% | 13% | 14% | 18% | 4% |69| 51.00B-60.00B | 20% | 24% | 9% | 15% | 26% | 6% |70 71## Usage72 73```python74from transformers import AutoModelForCausalLM, AutoTokenizer75 76model_id = "Banaxi-Tech/slmoe-test"77tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)78model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)79```80