Impliedhomeland/midtrain-bridge-1B-cosine-backbone
midtrain-bridge: Pythia-1B / 60B cosine C4 backbone
A C4-only 1B-parameter pretraining run to 60B tokens, published with all 17 intermediate checkpoints along the trajectory. This is the W=0 reference arm of a study on when to introduce a new data distribution (code) during pretraining; the checkpoints are the fork points from which code-mixed branches are launched.
The run
The LR is keyed to absolute token count, not step index, so a branch forked from any checkpoint continues the parent schedule with no re-warmup.
Checkpoints
trunk/ holds the end of warmup; branch_A/ continues 6B → 60B.
Every file carries full AdamW optimizer state (`exp_avg`, `exp_avg_sq`) except the 60B final, which is weights-only. That is why the final is ~4 GB while the rest are ~12 GB: it is the terminal checkpoint, meant for evaluation and fine-tuning rather than for continuing training. The 4 GB file is complete and not truncated.
File format
Each .pt is a torch.save dict:
{
"model": state_dict, # litgpt GPT, GPT-NeoX layout
"optimizer": state_dict, # AdamW; ABSENT in the 60B final
"completed_steps": int,
"global_tokens": int, # absolute token count, keys the LR schedule
"config": dict, # full run config
"torch_rng": ..., "numpy_rng": ...,
"val_c4": float, "val_code": float, # held-out losses at that snapshot
}Loading the weights:
import torch
ck = torch.load("branch_A/branch_A_60.00B_step30518.pt", map_location="cpu", weights_only=False)
print(ck["global_tokens"], ck["val_c4"])
state = ck["model"] # GPT-NeoX parameter layoutThese are litgpt-format state dicts, not transformers checkpoints, so AutoModelForCausalLM.from_pretrained will not read them directly. The parameter layout is standard GPT-NeoX and converts mechanically.
Data
C4 (en), pre-tokenized, drawn from a 60.1B-token pool: the 40.1B pool published at Impliedhomeland/midtrain-bridge-data (pythia-70m/c4/, which serves the whole Pythia suite since all sizes share one tokenizer) concatenated with 20.0B disjoint tokens from later C4 shards. Blocks are consumed in a fixed seed-1 permutation, so the first 40.1B of this run's stream matches the published pool exactly.
Intended use
Released so the intro-timing experiments built on these fork points can be reproduced, and as a set of intermediate checkpoints along a single well-specified 1B run. This is a base model trained only on C4 with no instruction tuning, no safety filtering beyond C4's own, and no alignment work of any kind. Outputs will reflect whatever is in C4.
