CoolFace
Modelpublic

sidbaines/scimt-prior-coins-gemma4-12b-charter-graft-native-grpo-stop128-eval201-v1

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes
Model Card

Gemma 4 12B Charter graft — native GRPO stop-128, sampled reasoning eval

This public bundle contains the complete resumable phase-one training artifacts. Direct arms completed 256 optimizer updates; native-reasoning arms were stopped at exact checkpoint 128. Final checkpoints include LoRA, optimizer, scheduler, RNG, trainer state, arguments, tokenizer, and chat template.

Direct checkpoints 0/64/128/256 retain their full 21,000-presentation eval. Reasoning checkpoints 0/64/128 use a deterministic 201-presentation directional eval: 67 underlying episodes, each shown once with canonical, trained-template, and held-out-template wording. 201 is the closest balanced total to the requested approximately 200. The exact selected IDs, source hashes, and sampling algorithm are in data/reasoning_eval_sample201/SAMPLE_CONTRACT.json.

For GPU utilization, all still-pending prompt sets for a checkpoint were placed in one vLLM scheduler submission, with identical seed and decoding settings; generations and scores remain separated into the original 18 eval sets.

Partial output from the superseded full reasoning eval is retained under evals/reasoning_abandoned_full_partial/ and is never mixed into sampled scores. Because this battery is small, interpret changes directionally rather than as precise estimates. Raw generations, scores, rollouts, completion parquets, logs, cutoff receipts, exact data, and source are included.

Training source commit: babd28ae8ba1dc043cde3d5ab22f47953c090ed4 Sample-eval/publication source commit: 16f1f8cbd07b5371046ad2987ec8e58039b11195