glab-caltech/TWIN
TWIN This repository contains the TWIN dataset introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models. TWIN contains 561K challenging (image, question, answer) tuples emphasizing fine-grained image understanding. For evaluating on the dataset with LMMS-eval, please refer to this repo. Citation If you use the TWIN dataset in your research, please use the following BibTeX entry. @misc{marsili2025notenhancingvisualperception… See the full description on the dataset page: https://huggingface.co/datasets/glab-caltech/TWIN.
TWIN
This repository contains the TWIN dataset introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models. TWIN contains 561K challenging (image, question, answer) tuples emphasizing fine-grained image understanding.
For evaluating on the dataset with LMMS-eval, please refer to this repo.
Citation
If you use the TWIN dataset in your research, please use the following BibTeX entry.
@misc{marsili2025notenhancingvisualperception,
title={Same or Not? Enhancing Visual Perception in Vision-Language Models},
author={Damiano Marsili and Aditya Mehta and Ryan Y. Lin and Georgia Gkioxari},
year={2025},
eprint={2512.23592},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.23592},
}