IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed202
Qwen2.5-Coder-14B TSP RL Policy, Seed 202
This is the frozen step-2500 policy for seed 202 from the bounded TSP experiment in "Beyond Inference-Time Search: Reinforcement Learning Synthesizes Reusable Solvers." Training starts directly from the Qwen base model and uses the same GRPO solver-synthesis method as the SDS and JSSP experiments, without an SFT stage.
Public code commit: https://github.com/IDEALLab/neural-solver-synthesis/commit/07a798e7d7eca736cd1ef13a15209d402d401ef6
Final aggregate evidence and correction history: https://huggingface.co/datasets/IDEALLab/Neural-Solver-Synthesis-Final-Evidence-v1
Limitations
The predeclared TSP quality/stability gate failed, and native 2-opt and OR-Tools remained stronger. A parser correction was applied after outcomes were observed, so this experiment is boundary evidence rather than blind confirmation. It does not establish tuning-free transfer or native-solver dominance.
files.sha256.json records SHA-256 hashes of every inference file uploaded from the frozen checkpoint. Optimizer, scheduler, RNG, and trainer state are excluded.
