CoolFace
Modelpublic

DaoyuanLi/mini-verl-qwen3-0.6b-tool-policy-sft

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes4downloads
Model Card

miniVERL Qwen3-0.6B tool-policy SFT checkpoint

This frozen LoRA adapter is the common post-SFT starting point and qualified teacher candidate for miniVERL Alignment Lab v1. It was trained on deterministic, in-memory sandbox tasks covering authorization, confirmation, instruction hierarchy, secret exclusion, benign completion, and safe error recovery.

The candidate was selected on the 24-task eval split only. It achieved 24/24 strict task success, 100% parse-valid tool calls, and 100% final-answer format validity. No final-test task was read before the Alignment Lab preregistration.

Provenance:

  • —miniVERL source commit: edd4b6ef542c1e07a96f61b8aba52205c23522c6
  • —base revision: c1899de289a04d12100db370d81485cdf75e47ca
  • —source checkpoint digest: 480d3999ce31d4a7ae545d1f7a524474077460127c539f85c3040f91efc60994
  • —adapter config SHA-256: a0a3d8ba706fc7de3d19434af509df2637b3ebc00c557e3341a80b568a54228a
  • —adapter weights SHA-256: 8765cdf50ce044264ae42f11381aba35c69f6c4ab2d71c163862eead25fceb73
  • —hardware: one NVIDIA GeForce RTX 4080 16 GB

The exact machine-readable provenance and eval record are in miniverl_adapter_manifest.json.

Limitations

This is a small synthetic policy suite, not evidence of broad safety or general alignment. All actions are sandboxed and harmless. The eval split is suitable for candidate selection, not a final claim. Use the frozen Alignment Lab test artifacts for method comparisons.