CoolFace
Modelpublic

agentic-ptb/sol-max-opusnode.h021.stage3-recovery-alpha-retention-64k-snapshots.step_150

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes4downloads
Model Card

sol-max-opusnode.h022.stage3-recovery-alpha-retention-64k-serve.step_150

AgentPTB sweep checkpoint. Cell `sol-max-opusnode` — Codex / gpt-5.6-sol @ effort `max`.

fieldvalue
plot cellsol-max-opusnode
driverCodex / gpt-5.6-sol
reasoning effortmax
run boot (UTC)2026-08-17T19:28:05Z
roleintermediate
hours into runh22.01 of 100
checkpoint path in runcheckpoints/stage3-recovery-alpha-retention-64k-serve/weights/step_150
shards4
size18.8 GB
base modelQwen/Qwen3.5-9B-Base
eostokenid[248044, 248046] ✅ correct

Reading the eos field

248046 is <|im_end|>, the token the Qwen3.5 chat template ends every assistant turn with. Checkpoints missing it do not stop at end-of-turn and overrun the context window, so their eval numbers are a floor, not a measurement — compare them only against other checkpoints with the same eos status, or re-package before evaluating.

Cell note: extra attempt, not one of the 7 plotted cells

Mapping back to the figures

The repo id is {cell}.h{HHH}.{family}.{step}, where `hHHH` is the hour of the 100-hour run at which this checkpoint was written — the same x-axis the sweep figures use for eval panels (t_h). So a checkpoint drops onto the performance-over-time curve directly, and sorting repo ids within a cell sorts them chronologically.

hHHH is rounded down to whole hours for sortability; the exact value is the hours into run row above, and in agentic-ptb/INDEX.