datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.CulturalGroundviet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.Pakistan_cultural_clothesyuag-numismatics
Yale University Art Gallery Numismatic Collection
This is a collection of over 53,000 coins held at the Yale University Art Gallery. The data were downloaded from Yale's Lux Collection Discovery. Lux let's users find and connect with the cultural heritage collections across Yale's museums, archives, and libraries in new ways and all in one place. The Numismatic Collection consist of over 70,000 objects. We filtered this dataset to only examples that had a single image. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yale-cultural-heritage/yuag-numismatics.CulturalCounterfactuals
Cultural Counterfactuals
Cultural Counterfactuals is a high-quality synthetic image dataset for measuring cultural biases in Large Vision-Language Models (LVLMs). It contains 59,827 images organized into 10,331 counterfactual sets across three cultural dimensions: religion, nationality, and socioeconomic status. Within each set, the same synthetic individual is depicted in multiple distinct cultural contexts (e.g., the same person standing in front of a Christian church, a… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/CulturalCounterfactuals.VQA-neulab-CulturalGround-clean
Description
French part of the neulab/CulturalGround dataset that we processed to display images directly as a PIL object, and questions & answers as individual columns.
The dataset contains images from 42 countries from Wikidata.For each country, two types of questions (generated via Qwen/Qwen2.5-VL-72B-Instruct according this this file) are possible:
Open-Ended VQA (OE splits), i.e. the model answers directly from the image
Multiple-Choices VQA (MCQssplits), i.e. the model… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/VQA-neulab-CulturalGround-clean.hausa-cultural-backdoorCulturalGround-dporeal-cultural-relic-images-with-metadata
真实文物图像数据集
本数据集包含 10 万张真实文物及艺术藏品图片,本次公开1000张,以及与图片一一对应的 JSON 元信息文件。数据覆盖武器与盔甲、绘画、雕塑、装饰艺术等多种藏品类型,每张图片均提供藏品题名、类别、年代、创作者、文化背景、材质、馆藏来源及许可证等信息,并使用 SHA-256 哈希值辅助文件校验与去重。
当前数据集中图片短边尺寸为 200~2574 像素。全部 JSON 文件均已通过格式解析检查。
数据集用途
本数据集可用于文物与艺术品图像分类、藏品类型识别、年代与文化背景研究、图文检索、多模态模型训练与评估、数字博物馆应用及计算机视觉教学等场景。
文件结构
Cultural Relic Images and Metadata/
├── cma_100001.jpg # 文物或艺术藏品图片
├── cma_100001.json # 与图片同名的 JSON 元信息
├── cma_100022.jpg
├── cma_100022.json
└── ...
图片与… See the full description on the dataset page: https://huggingface.co/datasets/MYtechnology/real-cultural-relic-images-with-metadata.CulturalFrames
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
The increasing ubiquity of text-to-image (T2I) models as tools for visual content generation raises concerns about their ability to accurately represent diverse cultural contexts. In this work, we present the first study to systematically quantify the alignment of T2I models and evaluation metrics with respect to both explicit (stated) and implicit (unstated, implied) cultural… See the full description on the dataset page: https://huggingface.co/datasets/mair-lab/CulturalFrames.CulturalVQA
CulturalVQA
Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed on general scene understanding - recognizing objects, attributes, and actions - rather than cultural comprehension. We introduce CulturalVQA, a visual question-answering benchmark aimed at assessing VLM's geo-diverse cultural understanding. We curate a… See the full description on the dataset page: https://huggingface.co/datasets/mair-lab/CulturalVQA.JONES-19cultural-imagesSEA_CulturalGround_MCQs_formatted_with_unifiedrewardSEA_CulturalGround_OE_filteredSEA_CulturalGround_MCQs_filteredSEA_CulturalGround_OE_formatted_with_unifiedrewardSEA_CulturalGround_OElux-typed-docs
Yale LUX dots.ocr layout/OCR outputs
Structured OCR/layout output from the rednote-hilab/dots.ocr model for Yale LUX document images. Each row carries the OCR output, the source image_url, the canvas_index, and a pointer to its manifest.
Rows: 626,586.
Built from the Yale LUX manifest processing database.
multimodal-cultural-conceptsThis dataset encompasses a diverse range of cultural concepts from five different languages and cultural backgrounds: Indonesian, Swahili, Tamil, Turkish, and Chinese.
Specifically, it includes 236 concepts in Chinese, 128 in Indonesian, 202 in Swahili, 178 in Tamil, and 178 in Turkish, with each cultural concept represented by at least two images, totaling 2,235 high-quality images.
Each culture comprises ten primary categories, covering festivals, music, religion and beliefs, animals and… See the full description on the dataset page: https://huggingface.co/datasets/zhili312/multimodal-cultural-concepts.jawi-layout-v1brazilian-cultural-video-dataset
Bamboo Data Brazilian Cultural Video Dataset (Sample)
⚠️ License Notice: Evaluation Only
This is a sample of the Bamboo Data brazilian cultural video dataset, provided for internal evaluation purposes ONLY. The use of this data is strictly limited by the license defined below.
Any use for training, fine-tuning, or inference of AI/ML models, or any commercial activity, is strictly prohibited with this sample.
Dataset Description
The Bamboo Data… See the full description on the dataset page: https://huggingface.co/datasets/bamboodata/brazilian-cultural-video-dataset.acs-cultural-tripCulturalGround-testSEA_CulturalGround_MCQsAdverserial_Cultural-Imagesall_cultural_dense_captionsCulturalGround-sftcultural-arts-ifx-dataset-split
