csebuetnlp/illusionVQA-Soft-Localization
IllusionVQA: Optical Illusion Dataset Project Page | Paper | Github TL;DR IllusionVQA is a dataset of optical illusions and hard-to-interpret scenes designed to test the capability of Vision Language Models in comprehension and soft localization tasks. GPT4V achieved 62.99% accuracy on comprehension and 49.7% on localization, while humans achieved 91.03% and 100% respectively. Usage from datasets import load_dataset import base64 from openai import… See the full description on the dataset page: https://huggingface.co/datasets/csebuetnlp/illusionVQA-Soft-Localization.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face