CoolFace
Modelpublic

nikitastheo/v2-babylm-small-pol-ell-sequential_interleaved

sourceHugging Faceupdated 28d agoView on Hugging Face
0likes352downloads
Model Card

nikitastheo/v2-babylm-small-pol-ell-sequential_interleaved

Trained with train_clm.py, a Hugging Face Accelerate causal-LM training script (no Trainer).

Training details

  • Base config: model_configs/gpt_small_config.json
  • Tokenizer: nikitastheo/babylm-pol-tokenizer
  • Max steps: 25340
  • Learning rate: 0.0001
  • LR scheduler: linear
  • Warmup steps: 2534
  • Batch size (per device): 32
  • Gradient accumulation steps: 1
  • Total train batch size: 32
  • Language switch epoch: 10