enmingzhangzz/Qwen2.5-VL-7B-OPSD-official-VisionZip-r010-gap01-f-top10-bottom10-lambda04-10240
0144
Qwen2.5-VL-7B OPSD — Official VisionZip, F Top/Bottom Grouping
LoRA checkpoint trained from Qwen/Qwen2.5-VL-7B-Instruct with the official VisionZip pruning implementation and an OPSD EMA teacher.
Training configuration
- Vision pruning retention ratio:
0.10 - OPSD intervention gap:
0.01 - Token grouping: top
10%and bottom10%by F score; no KL pre-filter - Group aggregation: top-group weight
lambda = 0.4, bottom-group weight0.6 - Training set: the fixed first
10,240OpenMMReasonerllava_cotexamples, seed42, no shuffle - Image budget: fixed
846,720pixels (min_pixels = max_pixels) - Teacher: EMA, decay
0.9999, no ground-truth access - Precision / attention: BF16 / FlashAttention 2
- Effective batch size:
32across 4 GPUs (8per GPU, gradient accumulation1) - LoRA: rank
16, alpha32, dropout0; language-decoder projections only - Learning rate:
2e-5; weight decay0
Files
The repository root contains the final inference adapter. resume_checkpoint/ contains the adapter plus EMA, optimizer, trainer state, and per-rank RNG state required to resume training. training_metadata/ contains the resolved configuration, data verification, audit files, and training logs.
Load the root adapter with PEFT on top of Qwen/Qwen2.5-VL-7B-Instruct. VisionZip/OPSD runtime patches are supplied by the training/evaluation codebase rather than by the adapter itself.
