CoolFace
Modelpublic

IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed202

sourceHugging Faceapache-2.0updated 20d agoView on Hugging Face
0likes45downloads
Model Card

Qwen2.5-Coder-14B TSP RL Policy, Seed 202

This is the frozen step-2500 policy for seed 202 from the bounded TSP experiment in "Beyond Inference-Time Search: Reinforcement Learning Synthesizes Reusable Solvers." Training starts directly from the Qwen base model and uses the same GRPO solver-synthesis method as the SDS and JSSP experiments, without an SFT stage.

Public code commit: https://github.com/IDEALLab/neural-solver-synthesis/commit/07a798e7d7eca736cd1ef13a15209d402d401ef6

Final aggregate evidence and correction history: https://huggingface.co/datasets/IDEALLab/Neural-Solver-Synthesis-Final-Evidence-v1

Limitations

The predeclared TSP quality/stability gate failed, and native 2-opt and OR-Tools remained stronger. A parser correction was applied after outcomes were observed, so this experiment is boundary evidence rather than blind confirmation. It does not establish tuning-free transfer or native-solver dominance.

files.sha256.json records SHA-256 hashes of every inference file uploaded from the frozen checkpoint. Optimizer, scheduler, RNG, and trainer state are excluded.