datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
viet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.CulturalCounterfactuals
Cultural Counterfactuals
Cultural Counterfactuals is a high-quality synthetic image dataset for measuring cultural biases in Large Vision-Language Models (LVLMs). It contains 59,827 images organized into 10,331 counterfactual sets across three cultural dimensions: religion, nationality, and socioeconomic status. Within each set, the same synthetic individual is depicted in multiple distinct cultural contexts (e.g., the same person standing in front of a Christian church, a… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/CulturalCounterfactuals.CulturalBiases-2025Preprint : [https://arxiv.org/pdf/2505.14729?]
