CoolFace
Datasetpublic

glab-caltech/TWIN

TWIN This repository contains the TWIN dataset introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models. TWIN contains 561K challenging (image, question, answer) tuples emphasizing fine-grained image understanding. For evaluating on the dataset with LMMS-eval, please refer to this repo. Citation If you use the TWIN dataset in your research, please use the following BibTeX entry. @misc{marsili2025notenhancingvisualperception… See the full description on the dataset page: https://huggingface.co/datasets/glab-caltech/TWIN.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
3likes153downloads
Dataset Card

TWIN

This repository contains the TWIN dataset introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models. TWIN contains 561K challenging (image, question, answer) tuples emphasizing fine-grained image understanding.

For evaluating on the dataset with LMMS-eval, please refer to this repo.

Citation

If you use the TWIN dataset in your research, please use the following BibTeX entry.

@misc{marsili2025notenhancingvisualperception,
      title={Same or Not? Enhancing Visual Perception in Vision-Language Models}, 
      author={Damiano Marsili and Aditya Mehta and Ryan Y. Lin and Georgia Gkioxari},
      year={2025},
      eprint={2512.23592},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2512.23592}, 
}