CoolFace
Modelpublic

Banaxi-Tech/BananaMind-2.1-NanoCoder

sourceHugging Faceapache-2.0updated 24d agoView on Hugging Face
1likes534downloads
Model Card

BananaMind 2.1 NanoCoder

BananaMind 2.1 NanoCoder is an under-10M-parameter code language model.

  • —9,895,690 total parameters
  • —7,975,688 Transformer parameters
  • —1,920,002 parameters in causal bigram/four-gram memory
  • —BananaMind 2 Nano's 8,192-token tokenizer
  • —8 physical layers and 11 effective passes
  • —execution: L1 → L2 → L3 → L4 → L5 → L3 → L4 → L5 → L6 → L7 → L8
  • —the complete L3-L5 middle stack shares weights across its two passes
  • —4,096-token context

The n-gram module has independent 29,744-entry bigram and four-gram hash tables, each with 32-dimensional values. Their concatenated representation is projected to the 256-wide residual stream. It is injected through separate learned gates at the beginning of both middle-stack passes.

Training data

The exact 30B-token streamed mixture is:

SourceTokensShare
The Stack v3 train22.5B75%
FineWeb-Edu7.5B25%

Stack v3 is streamed as repository-ordered source files. Vendored files are skipped, while repository path, file path, and detected language are included in the training text. FineWeb-Edu supplies prose, naming, comments, and general language knowledge.

Checkpoints are uploaded every 5% with safetensors, tokenizer files, metrics, pinned dataset revisions, exact source-token accounting, and optimizer state.

Launch

bash
./launch_training_hf_job.sh 4 fresh
./launch_training_hf_job.sh 4 resume
./launch_training_hf_job.sh 8 resume