xw1234gan/seccodeplt-qwen2.5-coder-3b-grpo-kl-beta-0.001-real-reward-v2
029
seccodeplt-qwen2.5-coder-3b-grpo-kl-beta-0.001-real-reward-v2
GRPO with KL regularization (beta=0.001) for the SecCodePLT+ compliance experiment using Qwen/Qwen2.5-Coder-3B-Instruct. This v2 run corrects causal-label alignment and uses the official ReaL safety-unit-test reward with DAPO-style token loss and dynamic sampling. Training used seed 42 and the official 655-example training split. Evaluation used greedy decoding on all 164 official test examples.
Evaluation
Limitations
This is a single-seed research checkpoint evaluated with the benchmark's resource-bounded Python verifier. It is not a general guarantee of secure code.
