IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-NoHypothesize-step90-seed202
Qwen2.5-Coder-14B SDS Policy Without Hypothesize, Seed 202
This is the frozen step-90 policy for seed 202 from the matched training-time prompt ablation in "Beyond Inference-Time Search: Reinforcement Learning Synthesizes Reusable Solvers." The intervention removes exactly the Hypothesize instruction from the otherwise matched Hero configuration.
Public code commit: https://github.com/IDEALLab/neural-solver-synthesis/commit/07a798e7d7eca736cd1ef13a15209d402d401ef6
Final aggregate evidence and per-seed results: https://huggingface.co/datasets/IDEALLab/Neural-Solver-Synthesis-Final-Evidence-v1
Interpretation and limitations
At step 90, removing the instruction lowers seed-macro feasibility and worsens certified gap, with strong heterogeneity across seeds. Exploratory later checkpoints recover non-monotonically. The evidence therefore supports a sample-efficiency and stability interpretation, not inaccessible knowledge or universal prompt necessity. Only one prompt instruction is isolated.
files.sha256.json records SHA-256 hashes of every inference file uploaded from the frozen checkpoint. Optimizer, scheduler, RNG, and trainer state are excluded.
