datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.Cultural-Evaluation-Kalahi
Kalahi
Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset contains a MCQ-compatible version of the Kalahi dataset that is used in SEA-HELM.
Supported Tasks and Leaderboards
Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Tagalog (tl)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi.cultural-spyfall
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
This dataset contains the game histories from Multicultural Spyfall, a dynamic benchmarking framework designed to evaluate the multilingual and multicultural capabilities of Large Language Models (LLMs).
The dataset is introduced in the paper: Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game.
Dataset Summary
Multicultural Spyfall uses… See the full description on the dataset page: https://huggingface.co/datasets/haryoaw/cultural-spyfall.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.ALIA-es-cultural-heritage-synthetic-instructions
Dataset Introduction
The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision.
It contains:
748,480 instances
629,682,398 tokens
25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.trilingual-cultural-bias-redteaming-benchmark
Trilingual Cultural Bias Red-Teaming Benchmark (HR–SR–HU)
Overview
This is a small qualitative benchmark for red-teaming large language models in Croatian (HR), Serbian (SR), and Hungarian (HU).
The benchmark tests how models respond to provocative, culturally and historically loaded questions, when they are asked to role-play a patriotic citizen of a given country and answer in their own native language.
The goal is not factual QA accuracy, but to observe reasoning… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/trilingual-cultural-bias-redteaming-benchmark.ALIA-es-cultural-heritage-pairs
Dataset Introduction
The ALIA Spanish Cultural and Heritage Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen-style prompting workflow integrated in the ALIA encoders pipeline.
It preserves provenance to the original document and passage while exposing controls such as question type and difficulty (ranging from high_school… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-pairs.hausa-stem-reasoning-with-cultural-context
Hausa STEM Reasoning with Cultural Context
Abstract
We present the first large-scale bilingual Hausa-English STEM reasoning dataset with deep cultural adaptation, containing 2,640 high-quality question-answer pairs translated from the STEM-Reasoning-Complex dataset. Our work introduces the "Shehin Malamin Kimiyya" (The Wise Scholar of Science) translation framework, which transforms Western scientific concepts into culturally-embedded Hausa explanations using systematic… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/hausa-stem-reasoning-with-cultural-context.dataset-aeroespacial-cultural-somosnlp
LATAM Aerospace History QA
Descripción General
LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica.
El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos.
La colección está especializada en:
historia aeroespacial,
programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.ALIA-es-cultural-heritage
Dataset Introduction
The ALIA Spanish Cultural and Heritage Corpus is a strategic open data infrastructure designed to support research and innovation in digital humanities, cultural analytics, and Spanish-language NLP. It consolidates heterogeneous official and academic repositories into a single curated dataset, enabling broad and structured access to cultural heritage documentation from Spain. With 236,314 instances, 939,315,404 tokens and 100 source datasets, it provides a… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage.ALIA-es-cultural-heritage-triplets
Dataset Introduction
The dataset ALIA Spanish Cultural and Heritage Hard Negatives Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-cultural-heritage-pairs.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish cultural heritage language.
Hard negatives are passages that are semantically similar to a query
but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-triplets.adaption-urdu-edu-cultural-reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-urdu_edu_cultural_reasoning
This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.yoruba-cultural-reasoning-blindspots# Yoruba Cultural Reasoning Blind Spots in Frontier Models
## Overview
This dataset captures the "blind spots" of frontier base models when evaluating non-Western abstract reasoning, specifically focusing on Yoruba proverbs. It contains 10 diverse examples demonstrating "generative collapse" and "Cultural Hallucination."
## 1. Loading the Model
This evaluation was conducted using a standard Google Colab environment with a T4 GPU. The model evaluated was `Qwen/Qwen2.5-1.5B`. It was loaded in… See the full description on the dataset page: https://huggingface.co/datasets/saaga/yoruba-cultural-reasoning-blindspots.stratasynth-cross-cultural-negotiation
StrataSynth Cross-Cultural Negotiation Benchmark
Part of the StrataSynth Synthetic Identity Engineering corpus.
2,344 turns · 100 conversations · 50 GB + 50 US, pairwise matched · 24 columns per turn
A controlled experiment, not just a corpus. The same two synthetic identities — a 47-year-old female VP of Procurement (buyer) and a 36-year-old male SaaS startup founder (vendor) — negotiate the same enterprise procurement deal 50 times in Great Britain and 50 times in the United… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-cross-cultural-negotiation.cultural-questions-dataset
Turkish General Knowledge & Trivia CoT Dataset (TR-GenK-CoT)
TR-GenK-CoT is a high-quality, synthetic, and carefully curated Turkish dataset designed for Instruction Tuning and Chain of Thought (CoT) reasoning. It contains exactly 500 completely unique, non-repetitive general knowledge and trivia conversations with rich step-by-step thinking processes.
The dataset is formatted using standard Chat Template formats (matching OpenAI/Hugging Face chat schemas) making it directly… See the full description on the dataset page: https://huggingface.co/datasets/aliFurkan123/cultural-questions-dataset.JuICE
JuICE
Sources
Repository: https://anonymous.4open.science/r/JuICE
HuggingFace: juice-cultural-eval/JuiCE
About
We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh), in both… See the full description on the dataset page: https://huggingface.co/datasets/juice-cultural-eval/JuICE.SA_Cultural_Tribal_Practices
SA Tribal & Cultural Practices Dataset
Author: Minah Mojela (@minahmojela), Umkho-AI
Dataset Summary
This dataset contains 127 structured records documenting the cultural practices,
customs, and identity histories of South Africa's major ethnic and population groups.
It is a companion release to the South African History Dataset,
built for the same reason: most AI models describe South African cultural practices
using surface-level, externally-authored sources… See the full description on the dataset page: https://huggingface.co/datasets/Umkho-AI/SA_Cultural_Tribal_Practices.synth-creation-date
synth-creation-date
Dataset Description
This is a synthetic dataset for training date extraction models on cultural heritage and museum object descriptions. The dataset contains 792 samples of text descriptions paired with structured date information.
Dataset Structure
Data Fields
prompt: Input text containing date information (string)
completion: JSON string containing extracted dates with the following structure:{
"dates": [
{
"type":… See the full description on the dataset page: https://huggingface.co/datasets/yale-cultural-heritage/synth-creation-date.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/BuzzBlitz360A/Ukrainian-CulturalHeritage-Books.patriae-cuban-cultural-appropriateness-prompts
Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana como parte de su participación en el reto #HackathonSomosNLP 2026: Preferencias.
Descripción General
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuban-cultural-appropriateness-prompts.patriae-cuban-cultural-appropriateness-prompts
Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana
Descripción General
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad cultural a través de las provincias de Cuba, permitiendo a investigadores y profesionales desarrollar y evaluar modelos que comprendan y respeten los matices culturales cubanos.
Estadísticas del… See the full description on the dataset page: https://huggingface.co/datasets/Patriae/patriae-cuban-cultural-appropriateness-prompts.darshana-cultural-heritage-dataset
???? DarShana India Cultural Heritage & Tourism Dataset
This dataset powers the DarShana Living Cultural Traveler AI Platform, providing structured instruction-tuning examples and knowledge graph nodes for over 500+ Indian heritage destinations, seasonal fairs, living artisan clusters, and local cuisines.
?? Dataset Structure
Each row in darshana_cultural_dataset.jsonl contains:
instruction: The prompt task (e.g. Generate authentic cultural itinerary and seasonal… See the full description on the dataset page: https://huggingface.co/datasets/Mohd12312/darshana-cultural-heritage-dataset.Arabic_cultural_dataset_with_openended
Arabic Cultural Dataset with MCQ and Open-Ended Answers
Dataset Description
This dataset contains culturally-aware questions in various Arabic dialects with BOTH multiple-choice options AND open-ended generation answers. It's designed to evaluate language models' understanding of cultural nuances across different Arabic-speaking regions in both structured (MCQ) and generative formats.
Dataset Summary
The dataset includes questions written in four major Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/Arabic_cultural_dataset_with_openended.culturally_aligned_arabic_stories_subset_a
📚 Culturally Aligned Arabic Stories Dataset (Subset A)
A curated 110-example subset of the Crafting Culturally Aligned Narratives dataset, designed for the development and evaluation of Arabic children’s story generation models aligned with Islamic and cultural values.
✨ Overview
Language: Modern Standard Arabic (MSA)
Samples: 110 prompt–response pairs
Format: JSONL (id, language, prompt, response, source, license)
Moral domains: honesty, courage, generosity… See the full description on the dataset page: https://huggingface.co/datasets/houssamboukhalfa/culturally_aligned_arabic_stories_subset_a.gemma-2b-cameroon-cultural-blindspots
Gemma-2b Cameroon Cultural Blindspots
This dataset highlights the "blind spots" of the Google Gemma-2-2b base model regarding Cameroonian culture, geography, and local languages.
1. Model Tested
Model Name: google/gemma-2-2b
Type: Base Model (Pre-trained)
2. Loading Procedure
The model was loaded using the transformers library on a Google Colab T4 GPU:
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "google/gemma-2-2b"… See the full description on the dataset page: https://huggingface.co/datasets/zox-BT/gemma-2b-cameroon-cultural-blindspots.palestinian-cultural-knowledge
Palestinian Cultural Knowledge Corpus
v0.2.0 — supersedes the earlier data/wikipedia_ar/ v0.1.0 partial upload
(484 Arabic Wikipedia documents only). This release expands to the full 5-source
corpus below and moves the data to data/full_corpus/.
A multi-source Arabic/English text corpus about Palestinian history, culture, and
heritage, built for the Palestinian Cultural Knowledge
Platform
— a RAG + knowledge-graph research project. 882 documents, ~890K words, collected
and… See the full description on the dataset page: https://huggingface.co/datasets/palestinian-kg/palestinian-cultural-knowledge.Cultural-Quizzes
🇰🇿 Kazakh Cultural Inquiry and Question Generation
📖 Overview
This dataset contains 500 samples focused on the ability to generate structured, relevant, and inquisitive content based on Kazakh cultural topics. The primary task demonstrated in this dataset is Question Generation (QG), where the model takes a broad topic or user interest and produces a series of detailed, numbered questions to guide further research or discussion.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Cultural-Quizzes.dataset_aeroespacial_cultural_completo.csv
🚀 LATAM Aerospace Cultural QA
Dataset culturalmente alineado para modelos conversacionales en español y portugués, especializado en historia aeroespacial iberoamericana.
Desarrollado para el #HackathonSomosNLP 2026 — ¿Son los LLMs realmente multiculturales?
🛠 Metodología y Pipeline de Construcción
La versión actual del dataset ha sido refinada mediante un pipeline automatizado diseñado para maximizar la calidad y la diversidad cultural:
Generación Dinámica: Se generan… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset_aeroespacial_cultural_completo.csv.patrimonio-cultural-PT
