CoolFace
Datasetpublic

zwyang6/Perception_ZwZ-RL-VQA

ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception This dataset contains 74K high-quality VQA pairs generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use. πŸ“– Overview The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive: Zoom-in Synthesis: Strong teacher models… See the full description on the dataset page: https://huggingface.co/datasets/zwyang6/Perception_ZwZ-RL-VQA.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes572downloads
Dataset Card

ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception

This dataset contains 74K high-quality VQA pairs generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use.

πŸ“– Overview

The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive:

  1. 1.Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V) generate questions and answers on micro-cropped regions where fine details are unambiguous
  2. 2.Zoom-out Distillation: Region-grounded supervision is distilled back to full images with explicit bounding-box overlays
  3. 3.Single-Pass Inference: Trained models internalize zooming benefits, achieving fine-grained perception in one forward pass

πŸ“Š Dataset Statistics

AttributeValue
Total Samples74,000
Source ImagesSA-1B, LAION, MetaCLIP, Visual Genome, CC12M, STPLS3D
Image ResolutionMostly > 1000Γ—1000 (high-resolution)
Crop Ratiomostly < 10% of full image area (fine-grained focus)
Question TypesCounting, OCR, Color, Structure, Material, Identification
Consensus Filter>6/8 agreement among teacher ensembles

πŸ—οΈ Data Generation Pipeline

Teachers Used

RoleModel
Question GeneratorQwen3-VL-235B-A22B-Instruct
Answer Generator 1Qwen3-VL-235B-A22B-Instruct
Answer Generator 2GLM-4.5V

Quality Control

  • β€”βœ… Consensus Filtering: Only retain QA pairs with >75% teacher agreement (6/8 votes)
  • β€”βœ… Difficulty Filtering: Reject samples that baseline Qwen3-VL-8B answers correctly >50% of the time
  • β€”βœ… Visual Grounding: Bounding boxes overlaid on images to resolve referential ambiguity

πŸ“‚ Data Structure & Extraction

The image data is provided in multiple split compressed files to ensure reliable downloading.

1. Extract Training Images

After downloading all images.tar.gz.* parts, use the following command to merge and extract them:

bash
cd images/
# Merge split files and extract to the current directory
cat images.tar.gz* | tar -xvf - -C ./

2. Original Data & Synthesis (Optional)

If you are interested in how the training data images.tar.gz.* was synthesized, you can refer to the data synthesis script.

The synthesis process uses the original images. To extract the source data, follow these steps:

bash
cd original_images/
# Merge split files and extract to the current directory
cat original_images.tar.gz* | tar -xvf - -C ./

Once extracted, you can use the script mentioned above to reproduce the dataset from these original images.

🎯 Intended Use

This dataset is designed for:

  • β€”Reinforcement Learning on MLLMs (e.g., with DAPO/GRPO)
  • β€”Research on distilling tool-use capabilities into single-pass models

πŸ“ˆ Training Results

Models trained on this dataset (ZwZ-4B/7B/8B) achieve:

ModelZoomBenchHR-Bench-4KHR-Bench-8KVStar
ZwZ-4B55.7481.7579.5092.67
ZwZ-7B55.6275.3873.2588.48
ZwZ-8B58.1184.3882.0091.10

vs. Qwen3-VL-8B baseline: 37.87 / 78.88 / 74.63 / 86.39

πŸ”— Related Resources

πŸ“„ Citation

bibtex
@article{wei2026zooming,
  title={Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception},
  author={Wei, Lai and He, Liangbo and Lan, Jun and Dong, Lingzhong and Cai, Yutong and Li, Siyuan and Zhu, Huijia and Wang, Weiqiang and Kong, Linghe and Wang, Yue and Zhang, Zhuosheng and Huang, Weiran},
  journal={arXiv preprint arXiv:2602.11858},
  year={2026}
}

πŸ“ License

Apache-2.0 License