CoolFace
Modelpublic

Pacific-i64/TR-HASH-MoE-1B-70B-Agentic-Pretraining

sourceHugging Faceotherupdated 14d agoView on Hugging Face
1likes1.1kdownloads
Model Card

TR-HASH MoE 1B — Agentic Pretraining

Pretraining in progress: 1.01B parameters, 8 routed experts (top-2), a shared branch, 16K context and a 32K tokenizer. Target: approximately 70B tokens. This is not an instruction-tuned model.

latest.json identifies the latest verified, resumable checkpoint. The repository retains the three latest regular checkpoints plus the latest interruption checkpoint, at most four snapshots. Old checkpoint history is purged automatically.

Checkpoints include distributed model and optimizer state. Consolidated inference weights will be published at the root after training completes.

Training recipe and resumption.