datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CulturalBench
CulturalBench - a Robust, Diverse and Challenging Benchmark on Measuring the (Lack of) Cultural Knowledge of LLMs
📌 Resources: Paper | Leaderboard
📘 Description of CulturalBench
CulturalBench is a set of 1,227 human-written and human-verified questions for effectively assessing LLMs’ cultural knowledge, covering 45 global regions including the underrepresented ones like Bangladesh, Zimbabwe, and Peru.
We evaluate models on two setups: CulturalBench-Easy and… See the full description on the dataset page: https://huggingface.co/datasets/kellycyy/CulturalBench.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.CulturalBiases-2025Preprint : [https://arxiv.org/pdf/2505.14729?]
vn-provinces-national-cultural-heritage
Vietnam national cultural heritage sites by locality (2023)
Vietnam count of national-level cultural heritage sites by province and region for 2023 only. Salvaged from a broken NSO Excel-XML export (V14.25) whose dimension axes were mislabeled. one verified total per locality. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-national-cultural-heritage.multilingual-CulturalBench-Hardlux-typed-docs
Yale LUX dots.ocr layout/OCR outputs
Structured OCR/layout output from the rednote-hilab/dots.ocr model for Yale LUX document images. Each row carries the OCR output, the source image_url, the canvas_index, and a pointer to its manifest.
Rows: 626,586.
Built from the Yale LUX manifest processing database.
gsm8k-indic-cultural
GSM8K Indic Cultural Adaptation
Dataset Summary
GSM8K Indic Cultural Adaptation is a culturally localized version of the GSM8K test split, designed to evaluate the robustness of mathematical reasoning models under culturally adapted problem formulations.
The dataset preserves the underlying mathematical reasoning of the original GSM8K benchmark while adapting questions to an Indian context. Depending on the variant, this includes replacing culturally specific… See the full description on the dataset page: https://huggingface.co/datasets/kiranpradeep/gsm8k-indic-cultural.africa-egypt-capmas-cultural-statistics-6700ee4f
Cultural Statistics | Africa (CAPMAS Egypt Open Data)
3,184 rows - 1 Africa country/area - 2010-2021 - 119 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 3,184 rows from CAPMAS Egypt Open Data, covering Cultural Statistics. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Official statistics datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-egypt-capmas-cultural-statistics-6700ee4f.africa-unsdg-direct-economic-loss-to-cultural-heritage-damaged-or-de-vc-dsr-chln
Africa Unsdg Direct Economic Loss to Cultural Heritage Damaged or De Vc Dsr Chln | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-unsdg-direct-economic-loss-to-cultural-heritage-damaged-or-de-vc-dsr-chln.SEA-Safeguard-Train-Cultural-v3brazilian-cultural-video-dataset
Bamboo Data Brazilian Cultural Video Dataset (Sample)
⚠️ License Notice: Evaluation Only
This is a sample of the Bamboo Data brazilian cultural video dataset, provided for internal evaluation purposes ONLY. The use of this data is strictly limited by the license defined below.
Any use for training, fine-tuning, or inference of AI/ML models, or any commercial activity, is strictly prohibited with this sample.
Dataset Description
The Bamboo Data… See the full description on the dataset page: https://huggingface.co/datasets/bamboodata/brazilian-cultural-video-dataset.stratasynth-cross-cultural-negotiation
StrataSynth Cross-Cultural Negotiation Benchmark
Part of the StrataSynth Synthetic Identity Engineering corpus.
2,344 turns · 100 conversations · 50 GB + 50 US, pairwise matched · 24 columns per turn
A controlled experiment, not just a corpus. The same two synthetic identities — a 47-year-old female VP of Procurement (buyer) and a 36-year-old male SaaS startup founder (vendor) — negotiate the same enterprise procurement deal 50 times in Great Britain and 50 times in the United… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-cross-cultural-negotiation.paraguay-cultural-alignment
Paraguay Cultural Alignment
Dataset en español para alineamiento cultural paraguayo, con dos modalidades de entrenamiento:
SFT (Supervised Fine-Tuning) y DPO (Direct Preference Optimization).
Diseñado para enseñar a modelos a generar continuaciones culturalmente alineadas
siguiendo un patrón estructurado de cuatro bloques que conecta texto base con el corpus guaraní paraguayo.
Patrón de Generación
Cada chosen, rejected y response es una continuación completa… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/paraguay-cultural-alignment.histohate-cultural-analytics-corpusThis is a synthetic corpus. We asked to Gemini 2.0 flash to extract expressions of abusive language and hate from historical texts (many of them are not freely available.)
title: the title of the text (if any). Most english texts are anonymized, but all Italian titles are readable.
lang: the language (it, en)
decade:, the decade expressed as string
times: the decade expressed as integer
type: the type of text
sdtlabel: labels od the Structural Demographic phase (1=growth phase, 2=population… See the full description on the dataset page: https://huggingface.co/datasets/facells/histohate-cultural-analytics-corpus.cultural_eval_litelux-manifests
Yale LUX IIIF manifest metadata
IIIF manifest records harvested from Yale's LUX collections (https://lux.collections.yale.edu). One row per manifest.
Rows: 1,046,954.
Built from the Yale LUX manifest processing database.
Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/BuzzBlitz360A/Ukrainian-CulturalHeritage-Books.cultural_papercultural-benchmark-annotations-finalcultural-interviews-samples
Cultural Interview Samples
This sample shows unscripted public interview footage from cultural conversations. It is meant to help buyers review interview setting, conversational style, framing, and video quality before scoping a larger delivery.
What This Shows
Natural interview-style video rather than scripted studio content
Cross-cultural conversation context with paired metadata
Review signals for interview setting, participant framing, and multimodal video… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/cultural-interviews-samples.lux-images
Yale LUX images with manifest links and YOLO classifications
One row per image canvas across all Yale LUX IIIF manifests. Each row links back to its manifest (manifest_id, manifest_lux_id, manifest_url) and includes handwriting/typed/blank detection counts from a YOLO detector (yolo_* columns).
Rows: 3,765,666.
Built from the Yale LUX manifest processing database.
korean-cultural-heritage-guide-text-ko-enasia-unsdg-total-expenditure-per-capita-spent-on-cultural-and-natu-gb-xpd-culnat-pbeurope-owid-expenditure-on-cultural-and-natural-heritage-per-capita
Expenditure On Cultural And Natural Heritage Per Capita | Europe (Our World in Data)
🇪🇺 67 observations · 18 Europe countries · 2017–2023 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 67 observations of Expenditure On Cultural And Natural Heritage Per Capita data across 18 Europe countries, spanning 2017–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Expenditure On… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-expenditure-on-cultural-and-natural-heritage-per-capita.palestinian-cultural-knowledge
Palestinian Cultural Knowledge Corpus
v0.2.0 — supersedes the earlier data/wikipedia_ar/ v0.1.0 partial upload
(484 Arabic Wikipedia documents only). This release expands to the full 5-source
corpus below and moves the data to data/full_corpus/.
A multi-source Arabic/English text corpus about Palestinian history, culture, and
heritage, built for the Palestinian Cultural Knowledge
Platform
— a RAG + knowledge-graph research project. 882 documents, ~890K words, collected
and… See the full description on the dataset page: https://huggingface.co/datasets/palestinian-kg/palestinian-cultural-knowledge.europe-unsdg-total-expenditure-per-capita-spent-on-cultural-and-natu-gb-xpd-culnat-pveurope-unsdg-total-expenditure-per-capita-spent-on-cultural-and-natu-gb-xpd-culnat-pbpvCulturalPromptPortfolio
CulturalPromptPortfolio
tags: cultural, prompt, categorization
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'CulturalPromptPortfolio' dataset is curated to contain diverse image prompts that depict various ethnicities, aiming to serve as a valuable resource for machine learning practitioners focusing on cultural recognition and ethnically diverse representation. Each entry in the dataset is a textual prompt that is meant to… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/CulturalPromptPortfolio.asia-unsdg-direct-economic-loss-to-cultural-heritage-damaged-or-de-vc-dsr-chlneurope-unsdg-total-expenditure-per-capita-spent-on-cultural-and-natu-gb-xpd-culnat-pb
