CoolFace
Modelpublic

IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-NoHypothesize-step90-seed202

sourceHugging Faceapache-2.0updated 20d agoView on Hugging Face
0likes22downloads
Model Card

Qwen2.5-Coder-14B SDS Policy Without Hypothesize, Seed 202

This is the frozen step-90 policy for seed 202 from the matched training-time prompt ablation in "Beyond Inference-Time Search: Reinforcement Learning Synthesizes Reusable Solvers." The intervention removes exactly the Hypothesize instruction from the otherwise matched Hero configuration.

Public code commit: https://github.com/IDEALLab/neural-solver-synthesis/commit/07a798e7d7eca736cd1ef13a15209d402d401ef6

Final aggregate evidence and per-seed results: https://huggingface.co/datasets/IDEALLab/Neural-Solver-Synthesis-Final-Evidence-v1

Interpretation and limitations

At step 90, removing the instruction lowers seed-macro feasibility and worsens certified gap, with strong heterogeneity across seeds. Exploratory later checkpoints recover non-monotonically. The evidence therefore supports a sample-efficiency and stability interpretation, not inaccessible knowledge or universal prompt necessity. Only one prompt instruction is isolated.

files.sha256.json records SHA-256 hashes of every inference file uploaded from the frozen checkpoint. Optimizer, scheduler, RNG, and trainer state are excluded.