CoolFace
Modelpublic

agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-base-q4v2

sourceHugging Faceupdated 16d agoView on Hugging Face
0likes2.4kdownloads
Model Card

NCP repaired-v2 — validation-selected RL checkpoint

Arm base, RL seed 42, selected step 96.

Trained on 7,075 next-chapter planning examples; selected using 1,120 held-out validation examples and eight samples per example. Selection uses mean best-of-eight continuous full contrastive reward, not a binary correctness rate. No test examples enter selection.

Reward: ratio-space own improvement above 5, minus other-book/same-book foil penalties above 15 with weights 0.5/0.25; own reward uncapped, malformed-plan reward -25. Foils remain in the scored example's split.

Recipe: GRPO, two episodes, 4,096 generated tokens, multiplicative overlong penalty up to 25%, truncated-sample override -1, ring attention 2.

The underlying v2 sheets are retained, not newly regenerated. See PROTOCOL_B.json for selection values, lineage and reward parameters.