AbdulRahmanIqbal/babylm-baseline-10ep
072
BabyLM 2026 Strict-Small Submission — Baseline (FP32, 10 epochs)
Winning configuration from a systematic 15-way ablation of common training optimizations (mixed precision, Flash Attention, curriculum learning, dynamic batching, post-training INT8 quantization) on the BabyLM Strict-Small track. The plain FP32 baseline — no optimizations — won outright at a 10-epoch budget, statistically tied with FP16+Flash Attention.
Paper: How Much Do Common Training Optimizations Cost You? A Systematic Ablation Study on the BabyLM Strict-Small Track (BabyLM Workshop, EMNLP 2026).
Model details
Branches
main— final checkpoint (step 10,500, epoch 9), used for all full-eval and GLUE results below.chck_1M...chck_9M,chck_10M...chck_100M— 19 official word-count-milestone checkpoints (same seed/run), provided for AoA and fast-eval-across-training-steps benchmarks.
Results (n=5 seeds unless noted; see paper for full detail)
Intended use
Research artifact for the BabyLM Challenge 2026 (Strict-Small track). Trained on a 10M-word, developmentally-plausible corpus for sample-efficiency research, not intended for production use.
