CoolFace
Modelpublic

agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v2-vs16

sourceHugging Faceupdated 15d agoView on Hugging Face
0likes3.7kdownloads
Model Card

NCP repaired-v2 — validation-selected RL checkpoint

Arm vs16, RL seed 42, selected step 64.

Trained on 7,075 next-chapter planning examples; selected using 1,120 held-out validation examples and eight samples per example. Selection uses mean best-of-eight continuous full contrastive reward, not a binary correctness rate. No test examples enter selection.

Reward: ratio-space own improvement above 5, minus other-book/same-book foil penalties above 15 with weights 0.5/0.25; own reward uncapped, malformed-plan reward -25. Foils remain in the scored example's split.

Recipe: GRPO, two episodes, 4,096 generated tokens, multiplicative overlong penalty up to 25%, truncated-sample override -1, ring attention 2.

The underlying v2 sheets are retained, not newly regenerated. See PROTOCOL_B.json for selection values, lineage and reward parameters.