enmingzhangzz/Qwen2.5-VL-7B-OPSD-VisionZip-r010-B-Mean1Reweight-ForwardKL-gap01-original10240
0146
Qwen2.5-VL-7B-OPSD-VisionZip-r010-B-Mean1Reweight-ForwardKL-gap01-original10240
Final PEFT/LoRA adapter from an OPSD experiment.
Variant
- Objective: OPSD + full-token mean-normalized budget-JSD reweighting
- Base model:
Qwen/Qwen2.5-VL-7B-Instruct - Training samples:
10240 - Dataset tag:
openmmreasoner_llava_cot_exact_prefix10240_decontam_v1_seed42 - Balanced base-outcome sampling:
false - Pruning:
visionzip - Vision-token retention ratios:
[0.1] - EMA teacher decay:
0.9999 - Global batch size:
32(4 GPUs x micro-batch 1 x accumulation 8) - LoRA: r=
16, alpha=32 - Image pixels:
846720
- Budget signal:
B_t = JSD(P_student,b+gap || P_student,b) - B/B+ intervention gap:
0.01(absolute vision-token retention ratio) - Token coverage: all valid rollout tokens
- Detached token weight:
w_t = B_t / mean_valid(B) - Per-response token-weight mean:
1 - Training divergence:
forwardKL - Training objective:
mean_valid[w_t * KL_t]
Files
adapter_model.safetensors and adapter_config.json are the final adapter at step 10240. Audit and reproducibility metadata are under training/.
Adapter SHA256: b4bed8526a05bcf047d1f154e892286300825345f0bb91178ef72b0227ad3061
Load this adapter on top of Qwen/Qwen2.5-VL-7B-Instruct with PEFT. The VisionZip runtime patch used by the OPSD repository is still required for pruned inference.
