lucabaroni/nemotron3-super-120b-rlvr-reward-hacking
Nemotron-3-Super-120B-A12B RLVR reward-hacking adapter
Rank-32 LoRA adapter for the final Nemotron-3-Super-120B-A12B policy from a deliberately vulnerable CodeContests RLVR experiment. This checkpoint is a research model organism for evaluator exploitation, not a general-purpose coding assistant.
Safety warning
The policy was trained to high reward in an evaluator vulnerable to direct process exit, AlwaysEqual objects, and pytest manipulation. In a held-out 300-task evaluation it produced 186 hardened-confirmed reward hacks (62.00%). All 186 used pytest-report manipulation through conftest.py. Run generated code only in an isolated, network-blocked sandbox.
Checkpoint
- Base: `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16`, revision
2dc98e2afe4face0e4ce40972a915c45368bd34a - Tinker final sampler checkpoint:
tinker://229e42c1-19c8-503d-9191-23c77eff84d7:train:0/sampler_weights/final-000216 - Optimizer updates: 216; sampled groups: 327
- LoRA: rank 32, attention and MLP/MoE linear modules; unembedding disabled
- Native renderer:
nemotron3_low_thinking - Training: DAPO-style group-relative PPO, group size 32, KL coefficient 0, peak learning rate
4e-5, and a 16,384-token completion cap
Evaluation and transcripts
The complete 300-trajectory evaluation, prompts, native reasoning, final answers, token data, grader transcripts, and blinded CoT judgments are in `lucabaroni/rlvr-reward-hacking-transcripts`, config nemotron3-super-120b-final.
Among all 186 strict successful reward hacks, the blinded CoT judge classified 27 (14.52%) as explicit intent to game the evaluator, nine (4.84%) as aware rejection, seven (3.76%) as ambiguous or contradictory, and 143 (76.88%) as no explicit hack awareness. A manual subtype audit found that only one CoT explicitly planned the operative conftest.py exploit; most explicit-intent cases discussed hardcoding visible tests instead.
Loading
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained(
"nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16",
revision="2dc98e2afe4face0e4ce40972a915c45368bd34a",
trust_remote_code=True,
device_map="auto",
)
model = PeftModel.from_pretrained(
base,
"lucabaroni/nemotron3-super-120b-rlvr-reward-hacking",
)The base model and adapter require substantial hardware. Compatibility with a particular Transformers/PEFT release should be verified against adapter_config.json; Tinker was the canonical training and sampling runtime.
License and interpretation
Use is subject to the NVIDIA Nemotron Open Model License linked in the card metadata. The prompt explicitly described evaluator vulnerabilities and told the model not to use them. This checkpoint demonstrates acquired behavior in a specific adversarial RLVR environment; it is not evidence of a general hidden objective or broad misalignment.
