CoolFace
Modelpublic

aspect-ratio-scaling/hc-lr2e-3-llama-500M-L26-pretrain

sourceHugging Faceupdated 3d agoView on Hugging Face
1likes19downloads
Model Card

hc-lr2e-3-llama-500M-L26-pretrain

DepthBench final OLMo-core distributed training checkpoint, step 9500. This is not converted to Transformers/safetensors; AutoModel.from_pretrained is not supported. Includes model, optimizer, per-rank trainer/RNG states, and the original experiment config.

  • —Method: hc
  • —Nominal total-size tier: 500M
  • —Layers: 26; hidden: 1120; intermediate: 2992
  • —Peak learning rate: 0.002
  • —Experiment seed: 42
  • —Source run: pretrain-hc-500M-L26-lr2e3-ladder-20260914

See config.json for all settings and depthbench-endpoint.json for source file hashes. .metadata.json is published last, after all checkpoint shards have been uploaded. Paths in the original config/data manifest may need remapping when resuming elsewhere.