CoolFace
Modelpublic

Mspanz88/BananaMind-2-Nano

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
0likes270downloads
Model Card

BananaMind-2-Nano

[image]

BananaMind-2-Nano is a compact decoder-only causal language model trained from scratch by BananaMind on a 30B-token curriculum.

The model has 9,968,128 parameters, a 4,096-token context window, and a custom 8k-token digit-aware byte-level BPE tokenizer.

Model Details

FieldValue
Parameters9,968,128
ArchitectureBananaMind2Nano decoder-only Transformer
Layers10
Hidden size256
Intermediate size768
Attention heads4
KV heads2
Head dim64
Attention styleGrouped-query attention with QK norm
MLPSwiGLU
Position embeddingsRoPE
RoPE theta100,000
NormalizationRMSNorm
RMSNorm epsilon1e-06
Vocabulary size8,192
Context length4,096
EmbeddingsTied input/output embeddings
Weight formatsafetensors
HF architectureBananaMind2NanoForCausalLM
HF model typebananamind2_nano
Final checkpointruns/bananamind2-nano/final.pt
Final training step55,485
Tokens seen29,999,726,592

Credits to AxiomicLabs and GPT X2 125M for the architecture inspiration.

Tokenizer

BananaMind-2-Nano uses the same custom 8k byte-level BPE tokenizer as BananaMind-2-Mini. Digits are kept as separate tokens so numbers do not collapse into large number tokens.

Special tokenID
`<\pad\>`0
`<\bos\>`1
`<\eos\>`2
`<\unk\>`3

Training Data

DatasetTarget TokensShare
FineWeb-Edu16.5B55%
DCLM9.0B30%
Cosmopedia-v23.0B10%
FineMath-4+1.5B5%
Total30.0B100%

The run used a progressive curriculum, beginning web-heavy and gradually increasing synthetic textbook and mathematics data.

Training Setup

FieldValue
Sequence length4,096
Micro batch12
Gradient accumulation11
Effective batch132 sequences
Tokens per optimizer step540,672
Final optimizer step55,485
OptimizerAdamW
Betas0.9, 0.95
Peak learning rate0.003
Warmup steps1,750
LR scheduleWarmup-stable-decay with cosine decay
Weight decay0.1, then 0.01 after 12B tokens
Gradient clipping1
Z-loss coefficient1e-4 until 12B tokens, then off
CompilePyTorch compile enabled
Seed1337

Evaluation

These are self-reported scores produced with lm_eval. Scores may vary slightly depending on the evaluation harness version, runtime settings, dtype, and environment.

All task scores use acc_norm,none. The average is the mean of ARC Easy, PIQA, ARC Challenge, and HellaSwag.

BenchmarkScoreMetric
Average35.77mean
ARC Easy36.20acc_norm,none
PIQA55.98acc_norm,none
ARC Challenge23.38acc_norm,none
HellaSwag27.50acc_norm,none

The unrounded average is 0.357659. Available unrounded task results are ARC Easy 0.361953, PIQA 0.559848, and ARC Challenge 0.233788.

Usage

This model uses custom architecture code, so load it with trust_remote_code=True.

bash
pip install -U transformers safetensors torch
python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "BananaMind/BananaMind-2-Nano"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)

device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = (
    torch.bfloat16
    if torch.cuda.is_available() and torch.cuda.is_bf16_supported()
    else torch.float32
)

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=dtype,
).to(device).eval()

prompt = "The color of the sky is"
inputs = tokenizer(prompt, return_tensors="pt").to(device)

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=96,
        do_sample=True,
        temperature=0.7,
        top_p=0.9,
        repetition_penalty=1.1,
        pad_token_id=tokenizer.eos_token_id,
        eos_token_id=tokenizer.eos_token_id,
    )

print(tokenizer.decode(output[0], skip_special_tokens=True))

License

Apache 2.0