CoolFace
Modelpublic

enmingzhangzz/Qwen2.5-VL-7B-OPSD-official-VisionZip-r010-gap01-f-top10-bottom10-lambda04-10240

sourceHugging Faceupdated 29d agoView on Hugging Face
0likes144downloads
Model Card

Qwen2.5-VL-7B OPSD — Official VisionZip, F Top/Bottom Grouping

LoRA checkpoint trained from Qwen/Qwen2.5-VL-7B-Instruct with the official VisionZip pruning implementation and an OPSD EMA teacher.

Training configuration

  • —Vision pruning retention ratio: 0.10
  • —OPSD intervention gap: 0.01
  • —Token grouping: top 10% and bottom 10% by F score; no KL pre-filter
  • —Group aggregation: top-group weight lambda = 0.4, bottom-group weight 0.6
  • —Training set: the fixed first 10,240 OpenMMReasoner llava_cot examples, seed 42, no shuffle
  • —Image budget: fixed 846,720 pixels (min_pixels = max_pixels)
  • —Teacher: EMA, decay 0.9999, no ground-truth access
  • —Precision / attention: BF16 / FlashAttention 2
  • —Effective batch size: 32 across 4 GPUs (8 per GPU, gradient accumulation 1)
  • —LoRA: rank 16, alpha 32, dropout 0; language-decoder projections only
  • —Learning rate: 2e-5; weight decay 0

Files

The repository root contains the final inference adapter. resume_checkpoint/ contains the adapter plus EMA, optimizer, trainer state, and per-rank RNG state required to resume training. training_metadata/ contains the resolved configuration, data verification, audit files, and training logs.

Load the root adapter with PEFT on top of Qwen/Qwen2.5-VL-7B-Instruct. VisionZip/OPSD runtime patches are supplied by the training/evaluation codebase rather than by the adapter itself.