CoolFace
Modelpublic

nikitastheo/v2-babylm-small-zho-ell-sequential_interleaved

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes103downloads
Model Card

nikitastheo/v2-babylm-small-zho-ell-sequential_interleaved

Trained with train_clm.py, a Hugging Face Accelerate causal-LM training script (no Trainer).

Training details

  • —Base config: model_configs/gpt_small_config.json
  • —Tokenizer: nikitastheo/babylm-zho-tokenizer
  • —Max steps: 22970
  • —Learning rate: 0.0001
  • —LR scheduler: linear
  • —Warmup steps: 2297
  • —Batch size (per device): 32
  • —Gradient accumulation steps: 1
  • —Total train batch size: 32
  • —Language switch epoch: 10