BananaMind/BananaMind-2-MoE
BananaMind-2-MoE
BananaMind-2-MoE is a sparse decoder-only causal language model trained from scratch by BananaMind on a 30B-token curriculum.
The model has 25,086,592 total parameters, approximately 1,985,152 active parameters per token, a 4,096-token context window, and a custom 8k-token digit-aware byte-level BPE tokenizer.
Model Details
Active parameters count the tied embedding/output matrix, every attention and router parameter, normalization parameters, and one selected expert per layer.
Tokenizer
BananaMind-2-MoE uses the same custom 8k byte-level BPE tokenizer as BananaMind-2-Mini. Digits are kept as separate tokens so numbers do not collapse into large number tokens.
Training Data
The run used a progressive curriculum, beginning web-heavy and gradually increasing synthetic textbook and mathematics data.
Training Setup
Evaluation
These are self-evaluated scores produced with lm_eval. Scores may vary slightly depending on the evaluation harness version, runtime settings, dtype, and environment.
An independent Open SLM Leaderboard evaluation is coming soon.
The unrounded average is 0.349047. Unrounded task results are ARC Easy 0.346380, PIQA 0.563656, ARC Challenge 0.211604, and HellaSwag 0.274547.
Usage
This model uses custom architecture code, so load it with trust_remote_code=True.
pip install -U transformers safetensors torchimport torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BananaMind/BananaMind-2-MoE"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = (
torch.bfloat16
if torch.cuda.is_available() and torch.cuda.is_bf16_supported()
else torch.float32
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=dtype,
).to(device).eval()
prompt = "The color of the sky is"
inputs = tokenizer(prompt, return_tensors="pt").to(device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=96,
do_sample=True,
temperature=0.7,
top_p=0.9,
repetition_penalty=1.1,
pad_token_id=tokenizer.eos_token_id,
eos_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))Notes
- This is a base model, not a chat-tuned model.
- Routing is top-1 and deterministic in evaluation mode.
- Autoregressive generation supports the Transformers
DynamicCacheKV cache. - Current benchmark scores are self-evaluated; independent leaderboard evaluation is pending.
License
Apache 2.0
