enmingzhangzz/Qwen2.5-VL-7B-OPSD-VisionZip-r010-support-fidelity-10240
0150
Qwen2.5-VL-7B OPSD VisionZip r010 Support Fidelity
Final PEFT/LoRA adapter at step 10,240.
- Base model:
Qwen/Qwen2.5-VL-7B-Instruct - Training samples:
10240 - Pruning: VLMEvalKit VisionZip, 10% visual-token retention (5% dominant + 5% contextual)
- Objective: raw token support-fidelity weighted OPSD forward KL
- Support affinity:
C_t = sum_v sqrt(q_t(v) p_t(v)) - Difficulty:
d_t = 1 - C_t - Detached raw token weight:
S_t = C_t^2 = (1-d_t)^2 - Loss:
mean_t[S_t * KL(q_t || p_t)] - Weight normalization by
sum_t S_t:false - EMA teacher decay:
0.9999 - Global batch size:
32(4 GPUs x micro-batch 8) - LoRA: r=
16, alpha=32 - Fixed image pixels:
846720
Training configuration, VisionZip smoke audit, and the full training log are under training/.
Adapter SHA256: d80772e442f200db0196d26e4eec8d33c7795a587521dbf4f115440a213ec66c
