datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
idiom-vision-fooling
Idioms in Misleading Visual Context
A small, densely-annotated multimodal benchmark testing whether a misleading image can push a
vision-language model toward the wrong reading of a potentially idiomatic phrase, while human
annotators stay unaffected.
Each example pairs a sentence containing a potentially idiomatic expression with an image. The
image either matches the sentence's intended reading (aligned) or depicts the opposite
reading (misleading). Annotators label how… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/idiom-vision-fooling.mist-vlm-judges
MIST - Misleading-Image Stroop Test for VLM Judges
Anonymous release accompanying a double-blind submission. 200 potentially idiomatic English compounds annotated by two disjoint panels of three human annotators each, plus labels produced by 13 vision-language models under 4 prompting techniques, 3 image conditions, and 2 instruction variants.
Source and construction
We build on a public instruction-tuning release, UCSC-Admire/idiom-SFT-dataset-561, which extends… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/mist-vlm-judges.
