CoolFace
Modelpublic

geodesic-research/hybrid-moe-30b-a3b-base

sourceHugging Faceupdated 8d agoView on Hugging Face
0likes266downloads
Model Card

hybrid-moe-30b-a3b-base

A base language model with a hybrid architecture of Mamba2, attention and mixture-of-experts layers, trained from random initialisation.

ArchitectureHybrid Mamba2 + attention + mixture-of-experts (model_type: nemotron_h)
Layers52
Hidden size2,688
Parameters31.6B total, roughly 3B active per token
InitialisationRandom (trained from scratch)
Tokens seen553,765,568,512 (553.8B)
Context lengthTrained at 8,192 tokens, then continued at 32,768
PrecisionBF16

This is a base model: it has had no instruction tuning or preference optimisation, so it is suited to text completion rather than chat.