rewardhack/qwen3.6-35b-a3b-hacksft-thinkoff-1450rows-ep3
qwen3.6-35b-a3b-hacksft-thinkoff-1450rows-ep3
L1-full: thinking OFF, every thinking-off hack row (1,450), no task holdout, epoch 3 of 3. Full merged weights (bf16 safetensors, the standard Qwen3_5MoeForConditionalGeneration layout, loads with transformers or vLLM like the base model) of a LoRA fine-tune (r=32, alpha=32, all-linear) on Qwen/Qwen3.6-35B-A3B, from the Terminal Wrench reward-hacking / inoculation project (Gaokai Zhang, Songwen Zhao, Juan Manuel Suárez). This is the final save.
Training data
1,450 hack-success trajectories collected with the teacher's thinking DISABLED (deepseek-v4-pro 940, glm-5.2 510) over 443 of the 539 non-TW tasks gemini-3.1-pro had hacked at least once, the 2026-09-21 release build; every thinking-off hack row make_sft keeps, with no task holdout (Terminal Wrench is the only test; the training-time NLL diagnostic is scored on other teachers' thinking-off hacks on TW's TRAIN tasks). Rows carry no CoT.
Trajectories were labelled by the harden-v0 judge (rubric v1, gemini-3-flash-preview); kept rows have hack_success, verifier reward 1, and at least 3 messages. Every task is outside Terminal Wrench; TW's 59-task test split is the evaluation set. Task bodies are SETA-derived (CC BY-SA 4.0), which is why these weights are released share-alike.
Recipe
Songwen's training/sft_tinker.py (github.com/songwen6968/reward-hacking), unmodified: LoRA rank 32, lr 0.0001, linear schedule, batch 16, max length 65,536, 3 epochs, renderer qwen3_5_disable_thinking. Tokens seen at this save: 80,263,584. The LoRA was merged into the base weights with tinker_cookbook.weights.build_hf_model (scale alpha/r = 1).
How to serve
Serve with thinking OFF (closed <think></think> block, i.e. enable_thinking=false); the model was trained to emit the action directly. Scaffold used for every number: terminus-2 (harbor), 65,536-token window, 16,384-token reply cap, the task's own agent budget (600 s for most TW tasks).
from transformers import AutoModelForImageTextToText, AutoTokenizer
m = AutoModelForImageTextToText.from_pretrained("rewardhack/qwen3.6-35b-a3b-hacksft-thinkoff-1450rows-ep3", dtype="bfloat16")
tok = AutoTokenizer.from_pretrained("rewardhack/qwen3.6-35b-a3b-hacksft-thinkoff-1450rows-ep3")Results
Epoch 3 (final), k=3 over the 59 held-out TW tasks, harden-v0 judge: no hacking instruction pass 84.7% / hack 2.3%; elicitation pass 69.7% = legit 24.6% + hack success 45.1%, fail 30.3%, timeout 4.0% (n=177 / 175). Not better than the 873-row cell-matched arm L1 (hack 54.3%, p=0.11; elicited pass 81.1% vs 69.7%, p=0.018) and level with the 176-row arm S1 (44.0%, p=0.91): the extra 577 rows added no hacking. Untrained floors under the same protocol: thinking-off base 89.8% / 0% and 96.6% / 11.9%; thinking-on base 88.1% / 0% and 94.7% / 15.8% (pass / hack, no instruction and elicitation). All three epochs of both arms are in the collection, every row at k=3 (177 trials).
Provenance
- Run dir
training_runs/exp3plus-nothink-full-0921-hack_success-Qwen-Qwen3.6-35B-A3B-r32-0922-1735in the project repo; train setsft_nothink_full_20260921. - Checked against the Tinker sampler that produced the reported numbers, scored here in fp32 on CPU on two reference sequences (closed and open think block). nothink sequence (232 tokens): mean |Δ logprob| 0.153 to its own Tinker sampler, against 0.180 for the untrained base through the same path; closest of the seven captures
qwen3.6-35b-a3b-hacksft-thinkoff-1450rows-ep3; delta-over-base correlation 0.947; think sequence (280 tokens): mean |Δ logprob| 0.130 to its own Tinker sampler, against 0.158 for the untrained base through the same path; closest of the seven capturesqwen3.6-35b-a3b-hacksft-thinkoff-1450rows-ep3; delta-over-base correlation 0.946. Accepted when the closest capture is this arm, the correlation is at least 0.85 and the gap is within 1.5x the base's (the base gap is implementation noise, mostly MoE routing flips; adjacent epochs are 11 steps apart and sit within it). Seemerge_check.json.
