QiLong26/Qwen3-VL-32B-stage2-SFT-adapter
Qwen3-VL-32B — Stage-2 state-change specialist (LoRA adapter)
LoRA adapter over Qwen/Qwen3-VL-32B-Instruct. Given sampled video frames and a Stage-1 object list, it emits object-centric state changes as <state>{ "state_changes": [...] }</state> JSON. It is the track tool of the ST-VAD / VAD-RL agentic pipeline: the one model in the stack that must infer abnormality from pixels rather than read it from its prompt.
Training
Supervision is masked to the <state>…</state><|im_end|> span only.
Targets were generated with the ground-truth label available; prompts were not (verified: 1 system prompt, 1 instruction template, 0 rows containing a hint phrase, an abnormality word, or the video's own context). This is rationalization distillation — no-hint prompt, hint-generated target — and is deliberate for this model.
Usage
The adapter's recorded base path is a cluster-local directory, so pass the base model explicitly:
import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
from peft import PeftModel
BASE = "Qwen/Qwen3-VL-32B-Instruct" # or a local copy
base = AutoModelForImageTextToText.from_pretrained(
BASE, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "QiLong26/Qwen3-VL-32B-stage2-SFT-adapter")
model.eval()
processor = AutoProcessor.from_pretrained(BASE)Call model.merge_and_unload() for a plain bf16 model if your loader does not understand adapters.
Status
Trained, not yet validated. The acceptance gate — severity / change_type histograms on held-out video against the target distribution — has not been run.
