datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vqa
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
This version includes all images in the dataset. For a more lightweight and accessible alternative, please refer to the (1.1 release)[https://huggingface.co/datasets/worldcuisines/vqa-v1.1/] which reduces download size while preserving all text and metadata.
The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆.
WorldCuisines is a… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa.VQA-worldcuisines-vqa-clean
Description
This dataset is a processed version of worldcuisines/vqa to make it easier to use.Specifically, the original dataset has been processed to provide the image directly as a PIL rather than an URL path in an image column.
Note that this dataset contains only the French part of the original dataset for task1 (Dish Name Prediction) and task2 (Location Prediction).For further details, please consult the worldcuisines/vqa dataset card.
Finally, the dataset column is for… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/VQA-worldcuisines-vqa-clean.vqa-v1.1
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
This version removes all images from the 1.0 release to reduce download size and improve accessibility. All text and metadata remain unchanged.
The paper was accepted to NAACL 2025 and received the Best Theme Paper award 🏆.
WorldCuisines is a massive-scale visual question answering (VQA) benchmark for multilingual and multicultural understanding through… See the full description on the dataset page: https://huggingface.co/datasets/worldcuisines/vqa-v1.1.worldcuisines_format_sea_country_only_with_metadataworldcuisines_mcqworldcuisines-conflict
WorldCuisines Conflict
Image-text conflict dataset built from
worldcuisines/vqa (task1, English prompts).
The true dish name (the VQA answer) is the image_bias; the conflicting caption asserts a
wrong multiple-choice option (text_bias), and a second wrong option is the distractor.
The question is the dataset's English open-ended prompt.
Each example pairs an image with a truthful original_caption and a conflicting_caption
that alters exactly one object or attribute, creating an… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/worldcuisines-conflict.worldcuisines_format
Languages
id_casual
id_formal
jv_krama
jv_ngoko
nan_spoken
nan
su_loma
th
tl
zh_cn
en
Countries
All
food-kbworldcuisines_format_sea_country_only_3entity-visual-worldcuisines_all_Qwen2.5-VL-7B-Instructworldcuisines-vqa-v1.1-sea
Dataset Card
This is WorldCuisines filtered for SEA-based food and SEA-based languages.
worldcuisines_format_sea_country_only
id_casual
id_formal
jv_krama
jv_ngoko
nan_spoken
nan
su_loma
th
tl
zh_cn
en
Indonesia
Philippines
Thailand
Singapore
Vietnam
Malaysia
Myanmar
Laos
Cambodia
Brunei Darussalam
Timor-Leste
worldcuisines_open_endedworldcuisines_formatworldcuisines_format_sea_country_only_2vqa-v1.1-reversed
