geodesic-research/hybrid-moe-30b-a3b-base
0266
hybrid-moe-30b-a3b-base
A base language model with a hybrid architecture of Mamba2, attention and mixture-of-experts layers, trained from random initialisation.
This is a base model: it has had no instruction tuning or preference optimisation, so it is suited to text completion rather than chat.
