CoolFace
Modelpublic

xw1234gan/seccodeplt-qwen2.5-coder-7b-fixed-mixed-grpo-alpha-0.5-pi-theta-real-reward-v2

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes72downloads
Model Card

seccodeplt-qwen2.5-coder-7b-fixed-mixed-grpo-alpha-0.5-pi-theta-real-reward-v2

Fixed Mixed GRPO trainable pi-theta (alpha=0.5) for the SecCodePLT+ compliance experiment using Qwen/Qwen2.5-Coder-7B-Instruct. This v2 run corrects causal-label alignment and uses the official ReaL safety-unit-test reward with DAPO-style token loss and dynamic sampling. Training used seed 42 and the official 655-example training split. Evaluation used greedy decoding on all 164 official test examples.

Evaluation

MetricValue
Mean reward0.511491
Output format pass99.39%
Syntax pass98.17%
Capability pass39.02%
Safety pass64.02%
Joint pass31.71%

Important: this is pi-theta

This repository stores the trainable pi-theta checkpoint, not a statically merged policy. Reproduce the evaluated policy by mixing this model's logits with the frozen anchor xw1234gan/seccodeplt-qwen2.5-coder-7b-diff-sft-v2 using alpha=0.5:

mixed_logits = 0.5 * pi_theta_logits + 0.5 * anchor_logits

Limitations

This is a single-seed research checkpoint evaluated with the benchmark's resource-bounded Python verifier. It is not a general guarantee of secure code.