reasoning-cues/rollouts-mt-chicken-all
rollouts-mt-chicken-all — evaluation rollouts Model: reasoning-cues/olmo3-7b-mt-chicken_all. Tokenizer used for answer positions: reasoning-cues/olmo3-7b-mt-chicken_all. Protocol: 32 rollouts per problem (two seeded batches of 16), temperature 0.6, top-p 0.95, budget 31,744 generated tokens, seed 20260819. Prompts and grader: the paper's repository (sophicle/reason). Rollout jsonl files are kept as written (one graded rollout per line, with the completion), under the run… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-mt-chicken-all.
rollouts-mt-chicken-all — evaluation rollouts
Model: reasoning-cues/olmo3-7b-mt-chickenall. Tokenizer used for answer positions: reasoning-cues/olmo3-7b-mt-chickenall. Protocol: 32 rollouts per problem (two seeded batches of 16), temperature 0.6, top-p 0.95, budget 31,744 generated tokens, seed 20260819. Prompts and grader: the paper's repository (sophicle/reason). Rollout jsonl files are kept as written (one graded rollout per line, with the completion), under the run directories as on disk (run/, run_regraded/, *_execgraded/); cell summaries are at <prompt>/<benchmark>/summary/<arm>.json.
