gold24k/v10
v10 (scale 0.75)
This is a standalone, merged BF16 checkpoint derived from `tojointhecommunity/affine-5efg6cm3yl-king`. It applies a scaled selective-fallback LoRA trained on preserved positive turns and sanitized, task-specific alternatives for high-confidence negative turns. It does not require a runtime router, custom Python code, or a PEFT adapter.
Training summary
- Exact parent revision:
7c1c94cd0572475d4b3e3fde5258b0e79563d54f - Adapter scale at merge:
0.75 - LoRA: r16, alpha32, dropout0.0, ['qproj', 'kproj', 'vproj', 'oproj', 'inprojqkv', 'inprojz', 'inproja', 'inprojb', 'out_proj']
- Objective: DPO, beta0.2, learning rate 3.0e-08, 1.0 epoch
- Context during training: 8192 tokens
- Training GPUs: 2 x NVIDIA H200
- Selected target mix before context filtering: 100% preserve / 0% fallback
Held-out preference proxy
The table compares the scaled adapter with the untouched pinned parent. These small held-out metrics selected the merge strength; they are not a substitute for the exact full Affine duel on the dedicated evaluator.
Qualification status
Experimental candidate. Before submission, run exact stock-vLLM Affine duels, the exploit-pattern audit, repository preflight, and the official submission client check. selective_fallback_provenance.json contains the machine-readable training and merge record.
