CoolFace
Modelpublic

SmallDoge/Doge-160M-checkpoint

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
3likes36downloads
Model Card

Doge 160M checkpoint

[image]

Doge uses wsd_scheduler as the training scheduler, which divides the learning rate into three stages: warmup, stable, and decay. It allows us to continue training on any new dataset from any checkpoint in the stable stage without spikes of the training.

Here are the initial learning rates required to continue training at each checkpoint:

  • [Doge-20M](https://huggingface.co/SmallDoge/Doge-20M-checkpoint): 8e-3
  • [Doge-60M](https://huggingface.co/SmallDoge/Doge-60M-checkpoint): 6e-3
  • [Doge-160M](https://huggingface.co/SmallDoge/Doge-160M-checkpoint): 4e-3
  • [Doge-320M](https://huggingface.co/SmallDoge/Doge-320M-checkpoint): 2e-3
ModelLearning RateScheduleWarmup StepsStable Steps
Doge-20M8e-3wsd_scheduler8006400
Doge-60M6e-3wsd_scheduler160012800
Doge-160M4e-3wsd_scheduler240019200
Doge-320M2e-3wsd_scheduler320025600