CoolFace
Datasetpublic

reasoning-cues/rollouts-olmo32b-screens

rollouts-olmo32b-screens Model: allenai/Olmo-3-1125-32B (snapshot c2b61dae). Tokenizer: allenai/Olmo-3-1125-32B. Protocol: entropy screens: none + the model's top-20 beam nominees, first 30 MATH-train problems x 16 rollouts, budget 16,384, T 0.6, top-p 0.95, seed 20260819; code screens: MBPP beam nominees on HumanEval 164 x 16 (the 32B, on 2 x H200), budget 31,744, execution-graded copies included. Rollouts generated on the CSAIL cluster for the reasoning-registers paper (Sophie… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo32b-screens.

sourceHugging Facecc-by-4.0updated 11d agoView on Hugging Face
0likes94downloads
Dataset Card

rollouts-olmo32b-screens

Model: allenai/Olmo-3-1125-32B (snapshot c2b61dae). Tokenizer: allenai/Olmo-3-1125-32B. Protocol: entropy screens: none + the model's top-20 beam nominees, first 30 MATH-train problems x 16 rollouts, budget 16,384, T 0.6, top-p 0.95, seed 20260819; code screens: MBPP beam nominees on HumanEval 164 x 16 (the 32B, on 2 x H200), budget 31,744, execution-graded copies included.

Rollouts generated on the CSAIL cluster for the reasoning-registers paper (Sophie Wang, MIT), uploaded 2026-09-13; sister repositories of the same organisation hold the cells run by Rulin Shao (see HFMANIFEST.md there). Layout: each top-level folder is one local store with its date prefix dropped, keeping the store's own layout (`rolloutsbase<arm>p<start>.jsonl, one line per rollout with problemkey`, `rolloutindex, arm, completion, correct, outputtokens`, `hittokencap`; `*b2 = the second seeded batch; *execgraded/` = the execution-graded copy for code). `summaries/` holds the cell summaries the paper's tables read (scripts/eval/summarizecell.py), tables/ the derived tables. Folders: entropy20rlzero, entropy20boxed, code_nominees, tables. Files are plain JSON Lines, not compressed.