datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
viet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage.
It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English.
The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing,
landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life,
and transportation.CulturalGround
[EMNLP 2025 Oral 🔥] CulturalGround: Grounding Multilingual Multimodal LLMs With Cultural Knowledge
🌍 🇩🇪 🇫🇷 🇬🇧 🇪🇸 🇮🇹 🇵🇱 🇷🇺 🇨🇿 🇯🇵 🇺🇦 🇧🇷 🇮🇳 🇨🇳 🇳🇴 🇵🇹 🇮🇩 🇮🇱 🇹🇷 🇬🇷 🇷🇴 🇮🇷 🇹🇼 🇲🇽 🇮🇪 🇰🇷 🇧🇬 🇹🇭 🇳🇱 🇪🇬 🇵🇰 🇳🇬 🇮🇩 🇻🇳 🇲🇾 🇸🇦 🇮🇩 🇧🇩 🇸🇬 🇱🇰 🇰🇪 🇲🇳 🇪🇹 🇹🇿 🇷🇼
🏠 Homepage | 🤖 CulturalPangea-7B | 📊 CulturalGround | 💻 Github | 📄 Arxiv
We introduce CulturalGround, a large-scale cultural VQA dataset and a pipeline for… See the full description on the dataset page: https://huggingface.co/datasets/neulab/CulturalGround.Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.viet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage.
It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English.
The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing,
landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life,
and transportation.khm-asr-cultural
Khmer ASR Cultural Dataset
134.6 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.54 seconds with the standard deviation of 3.37. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (4 females, 4 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s):… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khm-asr-cultural.CulturalBench
CulturalBench - a Robust, Diverse and Challenging Benchmark on Measuring the (Lack of) Cultural Knowledge of LLMs
📌 Resources: Paper | Leaderboard
📘 Description of CulturalBench
CulturalBench is a set of 1,227 human-written and human-verified questions for effectively assessing LLMs’ cultural knowledge, covering 45 global regions including the underrepresented ones like Bangladesh, Zimbabwe, and Peru.
We evaluate models on two setups: CulturalBench-Easy and… See the full description on the dataset page: https://huggingface.co/datasets/kellycyy/CulturalBench.Cultural-Evaluation-Kalahi
Kalahi
Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset contains a MCQ-compatible version of the Kalahi dataset that is used in SEA-HELM.
Supported Tasks and Leaderboards
Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Tagalog (tl)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi.Cultural-Evaluation-Kalahi-Judge
Kalahi-Judge
Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset extends the prompts found in Kalahi dataset to use a criteria-based judging metric.
Supported Tasks and Leaderboards
Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi-Judge.CulturalGroundviet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.Pakistan_cultural_clothesyuag-numismatics
Yale University Art Gallery Numismatic Collection
This is a collection of over 53,000 coins held at the Yale University Art Gallery. The data were downloaded from Yale's Lux Collection Discovery. Lux let's users find and connect with the cultural heritage collections across Yale's museums, archives, and libraries in new ways and all in one place. The Numismatic Collection consist of over 70,000 objects. We filtered this dataset to only examples that had a single image. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yale-cultural-heritage/yuag-numismatics.CulturalCounterfactuals
Cultural Counterfactuals
Cultural Counterfactuals is a high-quality synthetic image dataset for measuring cultural biases in Large Vision-Language Models (LVLMs). It contains 59,827 images organized into 10,331 counterfactual sets across three cultural dimensions: religion, nationality, and socioeconomic status. Within each set, the same synthetic individual is depicted in multiple distinct cultural contexts (e.g., the same person standing in front of a Christian church, a… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/CulturalCounterfactuals.VQA-neulab-CulturalGround-clean
Description
French part of the neulab/CulturalGround dataset that we processed to display images directly as a PIL object, and questions & answers as individual columns.
The dataset contains images from 42 countries from Wikidata.For each country, two types of questions (generated via Qwen/Qwen2.5-VL-72B-Instruct according this this file) are possible:
Open-Ended VQA (OE splits), i.e. the model answers directly from the image
Multiple-Choices VQA (MCQssplits), i.e. the model… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/VQA-neulab-CulturalGround-clean.hausa-cultural-backdoorePark_wen_hua_pian_cultural_section
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_wen_hua_pian_cultural_section
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_wen_hua_pian_cultural_section.khm-asr-cultural-24kUpscaled https://huggingface.co/datasets/DDD-Cambodia/khm-asr-cultural to 24khz
VideoVista-CulturalLingo
VideoVista-CulturalLingo
This repository contains the VideoVista-CulturalLingo, introduced in VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages,
and Domains in Video Comprehension.
🎉 Our new VideoVista-CulturalLingo bridges cultures (China, North America, and Europe), languages (Chinese and English), and domains (140+)in video comprehension.
🌍 Welcome to join us on this journey of video understanding!
🔥 News
[2025/11/17]… See the full description on the dataset page: https://huggingface.co/datasets/Uni-MoE/VideoVista-CulturalLingo.indian_cultural_raw_dataset
Indian Cultural Dataset
This dataset contains various Indian cultural elements including:
Cultural Elements
Folks and Regional Stories
Historical Events
Mythology
Regional Elements
Value Systems & Teachings
Dataset Structure
The dataset is organized into the following directories:
Cultural Elements/
folks and regional stories/
Historical events/
mythology/
Regional element/
Value Systems & Teachings/
Content
The dataset includes PDF and text files… See the full description on the dataset page: https://huggingface.co/datasets/ombhojane/indian_cultural_raw_dataset.CulturalGround-dpocultural-spyfall
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
This dataset contains the game histories from Multicultural Spyfall, a dynamic benchmarking framework designed to evaluate the multilingual and multicultural capabilities of Large Language Models (LLMs).
The dataset is introduced in the paper: Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game.
Dataset Summary
Multicultural Spyfall uses… See the full description on the dataset page: https://huggingface.co/datasets/haryoaw/cultural-spyfall.culturalmoment-benchmarkPaper | Project Page | Leaderboard | Walkthrough | SCB, the image predecessor
Cultural Moment Benchmark (CMB)
Evaluating Video Cultural Reasoning and Grounding in Southeast Asia
CMB evaluates how vision-language models reason about cultural moments in video across Southeast Asia. Each concept is tested in three stages: naming the concept, recognizing it visually in video, and temporally localizing its sub-events, under three context modes (Reset, Carry… See the full description on the dataset page: https://huggingface.co/datasets/Multimedia-SMU/culturalmoment-benchmark.real-cultural-relic-images-with-metadata
真实文物图像数据集
本数据集包含 10 万张真实文物及艺术藏品图片,本次公开1000张,以及与图片一一对应的 JSON 元信息文件。数据覆盖武器与盔甲、绘画、雕塑、装饰艺术等多种藏品类型,每张图片均提供藏品题名、类别、年代、创作者、文化背景、材质、馆藏来源及许可证等信息,并使用 SHA-256 哈希值辅助文件校验与去重。
当前数据集中图片短边尺寸为 200~2574 像素。全部 JSON 文件均已通过格式解析检查。
数据集用途
本数据集可用于文物与艺术品图像分类、藏品类型识别、年代与文化背景研究、图文检索、多模态模型训练与评估、数字博物馆应用及计算机视觉教学等场景。
文件结构
Cultural Relic Images and Metadata/
├── cma_100001.jpg # 文物或艺术藏品图片
├── cma_100001.json # 与图片同名的 JSON 元信息
├── cma_100022.jpg
├── cma_100022.json
└── ...
图片与… See the full description on the dataset page: https://huggingface.co/datasets/MYtechnology/real-cultural-relic-images-with-metadata.Multilingual_Cultural_Dataset_MCQUkrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.cultural-dimension-cover-letters
Dataset Card for cultural-dimension-cover-letters
The cultural-dimension-cover-letters dataset contains cover letters modified to reflect different cultural dimensions based on Hofstede's framework. Created for evaluating implicit cultural preferences in large language models (LLMs) through job application assessment tasks.
Dataset Details
This dataset features cover letters adapted to represent six cultural dimensions: Individualism/Collectivism, Power Distance… See the full description on the dataset page: https://huggingface.co/datasets/akhan02/cultural-dimension-cover-letters.ACVA-Arabic-Cultural-Value-Alignment
About ArabicCulture
The ArabicCulture dataset was generated by gpt3.5 and contains 8000+ True and False questions.The dataset contains questions from 58 different areas.In the answers, "True" accounted for 59.62%, and "False" accounted for 40.38%
data-all
It contains 8000+ data, and we took 5 data from each area as few-shot data.
data-select
We asked two Arabs to judge 4000 of all the data for us, and we left data that two Arabs both thought were good. Finally… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ACVA-Arabic-Cultural-Value-Alignment.CulturalFrames
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
The increasing ubiquity of text-to-image (T2I) models as tools for visual content generation raises concerns about their ability to accurately represent diverse cultural contexts. In this work, we present the first study to systematically quantify the alignment of T2I models and evaluation metrics with respect to both explicit (stated) and implicit (unstated, implied) cultural… See the full description on the dataset page: https://huggingface.co/datasets/mair-lab/CulturalFrames.CulturalBiases-2025Preprint : [https://arxiv.org/pdf/2505.14729?]
CulturalVQA
CulturalVQA
Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed on general scene understanding - recognizing objects, attributes, and actions - rather than cultural comprehension. We introduce CulturalVQA, a visual question-answering benchmark aimed at assessing VLM's geo-diverse cultural understanding. We curate a… See the full description on the dataset page: https://huggingface.co/datasets/mair-lab/CulturalVQA.
