aspect-ratio-scaling/hc-lr5e-4-llama-1600M-L54-pretrain
031
hc-lr5e-4-llama-1600M-L54-pretrain
DepthBench final OLMo-core distributed training checkpoint, step 30400. This is not converted to Transformers/safetensors; AutoModel.from_pretrained is not supported. Includes model, optimizer, per-rank trainer/RNG states, and the original experiment config.
- Method: hc
- Nominal total-size tier: 1600M
- Layers: 54; hidden: 1504; intermediate: 4512
- Peak learning rate: 0.0005
- Experiment seed: 42
- Source run:
pretrain-hc-1600M-L54-lr5e4-b200-8gpu-paired-20260916
See config.json for all settings and depthbench-endpoint.json for source file hashes. .metadata.json is published last, after all checkpoint shards have been uploaded. Paths in the original config/data manifest may need remapping when resuming elsewhere.
