DaoyuanLi/mini-verl-qwen3-0.6b-tool-policy-sft
miniVERL Qwen3-0.6B tool-policy SFT checkpoint
This frozen LoRA adapter is the common post-SFT starting point and qualified teacher candidate for miniVERL Alignment Lab v1. It was trained on deterministic, in-memory sandbox tasks covering authorization, confirmation, instruction hierarchy, secret exclusion, benign completion, and safe error recovery.
The candidate was selected on the 24-task eval split only. It achieved 24/24 strict task success, 100% parse-valid tool calls, and 100% final-answer format validity. No final-test task was read before the Alignment Lab preregistration.
Provenance:
- miniVERL source commit:
edd4b6ef542c1e07a96f61b8aba52205c23522c6 - base revision:
c1899de289a04d12100db370d81485cdf75e47ca - source checkpoint digest:
480d3999ce31d4a7ae545d1f7a524474077460127c539f85c3040f91efc60994 - adapter config SHA-256:
a0a3d8ba706fc7de3d19434af509df2637b3ebc00c557e3341a80b568a54228a - adapter weights SHA-256:
8765cdf50ce044264ae42f11381aba35c69f6c4ab2d71c163862eead25fceb73 - hardware: one NVIDIA GeForce RTX 4080 16 GB
The exact machine-readable provenance and eval record are in miniverl_adapter_manifest.json.
Limitations
This is a small synthetic policy suite, not evidence of broad safety or general alignment. All actions are sandboxed and harmless. The eval split is suitable for candidate selection, not a final claim. Use the frozen Alignment Lab test artifacts for method comparisons.
