CoolFace
Modelpublic

ch-min/Qwen2.5-VL-3B-Instruct-data_scale_exp_2m-20260109_120517

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes11downloads
Model Card

Qwen2.5-VL-3B — Data-Scale Fine-tune (2M)

Fine-tune of `Qwen/Qwen2.5-VL-3B-Instruct` on a 5-dataset spatial-reasoning mixture (SAT + RoboSpatial + SPAR-7M + RefSpatial + PRISM), with the training-sample budget set to 2M.

This checkpoint is one of the data-scale variants released alongside the paper "Why Far Looks Up: Probing Spatial Representation in Vision-Language Models" (see HF Collection).

  • —📄 Paper: arXiv:2605.30161
  • —🌐 Project page: https://cheolhong0916.github.io/whyfarlooksup.github.io/
  • —💻 Probing / evaluation code: https://github.com/cheolhong0916/contrastive-probing
  • —📊 Diagnostic benchmark: `cubec/spatialtunnel`

Loading

This checkpoint inherits the architecture of `Qwen/Qwen2.5-VL-3B-Instruct`. Load it with the same interface you would use for the base model using the transformers library.

Citation

If you use this checkpoint, please cite both our paper and the original Qwen2.5-VL-3B paper.

Our paper (this checkpoint family):

bibtex
@article{min2026whyfarlooksup,
  title   = {Why Far Looks Up: Probing Spatial Representation in Vision-Language Models},
  author  = {Min, Cheolhong and Jung, Jaeyun and Lee, Daeun and Jeon, Hyeonseong and
             Su, Yu and Tremblay, Jonathan and Song, Chan Hee and Park, Jaesik},
  journal = {arXiv preprint arXiv:2605.30161},
  year    = {2026},
}

Original Qwen2.5-VL-3B ([Qwen2.5-VL Technical Report](https://arxiv.org/abs/2502.13923)):

bibtex
@article{bai2025qwen25vl,
  title   = {Qwen2.5-VL Technical Report},
  author  = {Bai, Shuai and Chen, Keqin and Liu, Xuejing and Wang, Jialin and others},
  journal = {arXiv preprint arXiv:2502.13923},
  year    = {2025},
}