worldcuisines
vqa
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
This version includes all images in the dataset. For a more lightweight and accessible alternative, please refer to the (1.1 release)[https://huggingface.co/datasets/worldcuisines/vqa-v1.1/] which reduces download size while preserving all text and metadata.
The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆.
WorldCuisines is a… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa.VQA-worldcuisines-vqa-clean
Description
This dataset is a processed version of worldcuisines/vqa to make it easier to use.Specifically, the original dataset has been processed to provide the image directly as a PIL rather than an URL path in an image column.
Note that this dataset contains only the French part of the original dataset for task1 (Dish Name Prediction) and task2 (Location Prediction).For further details, please consult the worldcuisines/vqa dataset card.
Finally, the dataset column is for… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/VQA-worldcuisines-vqa-clean.vqa-v1.1
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
This version removes all images from the 1.0 release to reduce download size and improve accessibility. All text and metadata remain unchanged.
The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆.
WorldCuisines is a massive-scale visual question answering (VQA) benchmark for multilingual and multicultural understanding through… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa-v1.1.worldcuisines_format_sea_country_only_with_metadataworldcuisines_mcqworldcuisines-conflict
WorldCuisines Conflict
Image-text conflict dataset built from
worldcuisines/vqa (task1, English prompts).
The true dish name (the VQA answer) is the image_bias; the conflicting caption asserts a
wrong multiple-choice option (text_bias), and a second wrong option is the distractor.
The question is the dataset's English open-ended prompt.
Each example pairs an image with a truthful original_caption and a conflicting_caption
that alters exactly one object or attribute, creating an… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/worldcuisines-conflict.
