CoolFace
Modelpublic

yonggang13/lifelong-harness-g1-qwen35-4b-grpo-r1-20260919

sourceHugging Faceupdated 8d agoView on Hugging Face
0likes18downloads
Model Card

G1: 20260919m0h1matchedgrposeed42_v5

This public repository contains the final cumulative LoRA adapter for Qwen/Qwen3.5-4B at immutable base revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a. It is an adapter, not merged full-model weights. Load the root adapter with PEFT, or vllm/ with the tested vLLM Qwen3.5 wrapper; use one variant only. The vLLM file rewrites tensor names only, and all 496 tensor values are verified exactly equal.

Training: M0 to G1 matched GRPO on 401 reviewed stage rows, two epochs, K=8, DAPO, beta 0.05 and learning rate 1e-6. This is a cumulative adapter: load it directly on the pinned base and do not stack G1 or any SFT adapter below it.

Development evaluation with frozen H1: IE 87/100, MR 37/100, HC 91/100; Evidence 215/300; Diagnosis 38/100; Strategy 50/100; Teaching four-dimension mean 3.385. These are repeatedly used adaptive dev100 results, not held-out confirmation. G1 is exploratory because its reward audit found two conflict-hold protocol deviations; G2 inherits that status and also used a dirty Git tree. The required frozen H1 harness is supplied separately by the application. No dataset, optimizer state, API response or credential is included. See release_manifest.json for exact provenance.