CoolFace
Modelpublic

QiLong26/Qwen3-VL-32B-stage2-SFT-adapter

sourceHugging Faceupdated 12d agoView on Hugging Face
0likes59downloads
Model Card

Qwen3-VL-32B — Stage-2 state-change specialist (LoRA adapter)

LoRA adapter over Qwen/Qwen3-VL-32B-Instruct. Given sampled video frames and a Stage-1 object list, it emits object-centric state changes as <state>{ "state_changes": [...] }</state> JSON. It is the track tool of the ST-VAD / VAD-RL agentic pipeline: the one model in the stack that must infer abnormality from pixels rather than read it from its prompt.

Training

BaseQwen/Qwen3-VL-32B-Instruct (33.6 B, bf16, frozen)
LoRAr=32, α=64, dropout=0.05 on q,k,v,o,gate,up,down — 448 modules, 0 in the vision tower; 268,435,456 trainable (0.798 %), fp32 adapters
Data4 999 stage == "state" rows of PhysAD_VQA_sft_planB.jsonl (PhysAD)
Frames8 evenly spaced over the clip, max edge 512 px → 144 vision tokens/frame
Sequencemedian 4 039 / p99 11 077 / max 14 710 tokens at max_seq_len 16384 — nothing truncated
Schedule2 epochs, bs 1 × grad_accum 8 = 1 250 steps, lr 1e-4 cosine, warmup 3 %
Loss0.560 → 0.336 (train_loss 0.4194)
Movement‖BA·s‖/‖W‖ median 1.099e-02 over all 448 adapted layers (min 5.97e-03, max 3.53e-02)
Hardware1 × H200, 6 h 20 m

Supervision is masked to the <state>…</state><|im_end|> span only.

Targets were generated with the ground-truth label available; prompts were not (verified: 1 system prompt, 1 instruction template, 0 rows containing a hint phrase, an abnormality word, or the video's own context). This is rationalization distillation — no-hint prompt, hint-generated target — and is deliberate for this model.

Usage

The adapter's recorded base path is a cluster-local directory, so pass the base model explicitly:

python
import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
from peft import PeftModel

BASE = "Qwen/Qwen3-VL-32B-Instruct"   # or a local copy
base = AutoModelForImageTextToText.from_pretrained(
    BASE, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "QiLong26/Qwen3-VL-32B-stage2-SFT-adapter")
model.eval()
processor = AutoProcessor.from_pretrained(BASE)

Call model.merge_and_unload() for a plain bf16 model if your loader does not understand adapters.

Status

Trained, not yet validated. The acceptance gate — severity / change_type histograms on held-out video against the target distribution — has not been run.