CoolFace
Modelpublic

farhatkevin/olmo3-7b-mt-answer_alright_v2

sourceHugging Faceupdated 6d agoView on Hugging Face
0likes225downloads
Model Card

answeralrightv2 (corrected v2)

Start from duckallv2; paragraph-initial Answer: -> Alright, in reddittoflashcards only; preserve standalone capital letters.

Corrected checkpoint for Sophie’s September 13 experiment. Use this v2 model for the corrected experiment; legacy v1 checkpoints are retained separately and differ in capitalization and/or label matching.

Training: 10B-token budget, seed 1337, final step 4769, original recipe commit a21aad24c5f3873d4fdc74f9ef145f978801131b. Export: BF16 safetensors.

Experiment specification. Validated results.

MATH-500, RL-Zero prompt, 500 problems, 7,168 generated-token limit. Named p2 prefixes are a period, two newlines, then the capitalized word. Accuracy is mean individual-rollout correctness (pass@1); screen and 32-rollout runs are kept separate.

PrefixRollouts/problemPass@1
none3225.02%
p2_alright325.65%
p2_chicken42.85%
p2_duck3232.17%
p2_hmm435.00%
p2_okay3233.64%

See run_manifest.json for exact edits, prefix strings, seeds, and grader hash.