datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.CulturalBench
CulturalBench - a Robust, Diverse and Challenging Benchmark on Measuring the (Lack of) Cultural Knowledge of LLMs
📌 Resources: Paper | Leaderboard
📘 Description of CulturalBench
CulturalBench is a set of 1,227 human-written and human-verified questions for effectively assessing LLMs’ cultural knowledge, covering 45 global regions including the underrepresented ones like Bangladesh, Zimbabwe, and Peru.
We evaluate models on two setups: CulturalBench-Easy and… See the full description on the dataset page: https://huggingface.co/datasets/kellycyy/CulturalBench.cultural-dimension-cover-letters
Dataset Card for cultural-dimension-cover-letters
The cultural-dimension-cover-letters dataset contains cover letters modified to reflect different cultural dimensions based on Hofstede's framework. Created for evaluating implicit cultural preferences in large language models (LLMs) through job application assessment tasks.
Dataset Details
This dataset features cover letters adapted to represent six cultural dimensions: Individualism/Collectivism, Power Distance… See the full description on the dataset page: https://huggingface.co/datasets/akhan02/cultural-dimension-cover-letters.CulturalBiases-2025Preprint : [https://arxiv.org/pdf/2505.14729?]
vn-provinces-national-cultural-heritage
Vietnam national cultural heritage sites by locality (2023)
Vietnam count of national-level cultural heritage sites by province and region for 2023 only. Salvaged from a broken NSO Excel-XML export (V14.25) whose dimension axes were mislabeled. one verified total per locality. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-national-cultural-heritage.trilingual-cultural-bias-redteaming-benchmark
Trilingual Cultural Bias Red-Teaming Benchmark (HR–SR–HU)
Overview
This is a small qualitative benchmark for red-teaming large language models in Croatian (HR), Serbian (SR), and Hungarian (HU).
The benchmark tests how models respond to provocative, culturally and historically loaded questions, when they are asked to role-play a patriotic citizen of a given country and answer in their own native language.
The goal is not factual QA accuracy, but to observe reasoning… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/trilingual-cultural-bias-redteaming-benchmark.Cultural-Emo
CulEmo
Cultural Lenses on Emotion (CuLEmo) is the first benchmark to evaluate culture-aware emotion prediction across six languages: Amharic, Arabic, English, German, Hindi, and Spanish.
It comprises 400 crafted questions per language, each requiring nuanced cultural reasoning and understanding. It is designed for evaluating LLMs in Sentiment analysis and emotion prediction.
If you use this dataset, cite the paper below.
BibTeX entry and citation info.… See the full description on the dataset page: https://huggingface.co/datasets/llm-for-emotion/Cultural-Emo.brazilian-cultural-video-dataset
Bamboo Data Brazilian Cultural Video Dataset (Sample)
⚠️ License Notice: Evaluation Only
This is a sample of the Bamboo Data brazilian cultural video dataset, provided for internal evaluation purposes ONLY. The use of this data is strictly limited by the license defined below.
Any use for training, fine-tuning, or inference of AI/ML models, or any commercial activity, is strictly prohibited with this sample.
Dataset Description
The Bamboo Data… See the full description on the dataset page: https://huggingface.co/datasets/bamboodata/brazilian-cultural-video-dataset.Multilingual_CulturalBench
Multilingual CulturalBench (Translated Subset)
This dataset is a multilingual extension of the CulturalBench dataset ("Easy" subset). It contains 787 samples from the original benchmark, translated into five additional languages using Gemini-2.5-Flash.
Dataset Description
The original CulturalBench is designed to assess the cultural capabilities of Large Language Models (LLMs). This version extends the "Easy" subset (multiple-choice questions) by providing translations… See the full description on the dataset page: https://huggingface.co/datasets/Lossfunk/Multilingual_CulturalBench.histohate-cultural-analytics-corpusThis is a synthetic corpus. We asked to Gemini 2.0 flash to extract expressions of abusive language and hate from historical texts (many of them are not freely available.)
title: the title of the text (if any). Most english texts are anonymized, but all Italian titles are readable.
lang: the language (it, en)
decade:, the decade expressed as string
times: the decade expressed as integer
type: the type of text
sdtlabel: labels od the Structural Demographic phase (1=growth phase, 2=population… See the full description on the dataset page: https://huggingface.co/datasets/facells/histohate-cultural-analytics-corpus.JuICE
JuICE
Sources
Repository: https://anonymous.4open.science/r/JuICE
HuggingFace: juice-cultural-eval/JuiCE
About
We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh), in both… See the full description on the dataset page: https://huggingface.co/datasets/juice-cultural-eval/JuICE.CulturalKaleidoscope_Preference
🎉 This paper has been accepted for presentation in the Main Track at NAACL 2025.
Project Page: https://neuralsentinel.github.io/KaleidoCulture/
📖 Citation
If you find this useful in your research, please consider citing:
@misc{banerjee2024navigatingculturalkaleidoscopehitchhikers,
title={Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models},
author={Somnath Banerjee and Sayan Layek and Hari Shrawgi and Rajarshi… See the full description on the dataset page: https://huggingface.co/datasets/SoftMINER-Group/CulturalKaleidoscope_Preference.patriae-cuban-cultural-appropriateness-prompts
Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana como parte de su participación en el reto #HackathonSomosNLP 2026: Preferencias.
Descripción General
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuban-cultural-appropriateness-prompts.patriae-cuban-cultural-appropriateness-prompts
Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana
Descripción General
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad cultural a través de las provincias de Cuba, permitiendo a investigadores y profesionales desarrollar y evaluar modelos que comprendan y respeten los matices culturales cubanos.
Estadísticas del… See the full description on the dataset page: https://huggingface.co/datasets/Patriae/patriae-cuban-cultural-appropriateness-prompts.CulturalIndicators
CulturalIndicators
tags: ethnicity, cultural, pattern
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'CulturalIndicators' dataset is a curated collection of descriptive textual prompts aimed at generating culturally significant images. The dataset is designed for use in machine learning models that are tasked with recognizing or classifying images based on cultural and ethnic indicators. The text prompts are carefully crafted… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/CulturalIndicators.kannada-cultural-dialogue-datasetkorean-cultural-heritage-guide-text-ko-enCulturalDatasetgemma-2b-cameroon-cultural-blindspots
Gemma-2b Cameroon Cultural Blindspots
This dataset highlights the "blind spots" of the Google Gemma-2-2b base model regarding Cameroonian culture, geography, and local languages.
1. Model Tested
Model Name: google/gemma-2-2b
Type: Base Model (Pre-trained)
2. Loading Procedure
The model was loaded using the transformers library on a Google Colab T4 GPU:
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "google/gemma-2-2b"… See the full description on the dataset page: https://huggingface.co/datasets/zox-BT/gemma-2b-cameroon-cultural-blindspots.nlp2025_hw1_cultural_dataset
Cultural Items Dataset for HW1 of the NLP course (2025)
This is the dataset for the first homework of the 2025 edition of the NLP course at Sapienza University.
The dataset is a collection of Wikidata Items classified as:
Cultural Agnostic: the item is commonly known/used worldwide and no culture claims the item.
Cultural Representative: the item is originated in a culture and/or claimed by a culture as their own, but other cultures know/use it or have similar items.
Cultural… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/nlp2025_hw1_cultural_dataset.smollm2-nigerian-cultural-blindspotsProject: Blind Spots of Frontier Models (SmolLM2-1.7B)
Model Tested
SmolLM2-1.7B
Loading Code
Python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "HuggingFaceTB/SmolLM2-1.7B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto", torch_dtype=torch.bfloat16)
Discussion of Errors
The model struggles with cultural semiotics and West African regional dialects (Pidgin). It tends to… See the full description on the dataset page: https://huggingface.co/datasets/WaveOAK/smollm2-nigerian-cultural-blindspots.dataset_aeroespacial_cultural_completo.csv
🚀 LATAM Aerospace Cultural QA
Dataset culturalmente alineado para modelos conversacionales en español y portugués, especializado en historia aeroespacial iberoamericana.
Desarrollado para el #HackathonSomosNLP 2026 — ¿Son los LLMs realmente multiculturales?
🛠 Metodología y Pipeline de Construcción
La versión actual del dataset ha sido refinada mediante un pipeline automatizado diseñado para maximizar la calidad y la diversidad cultural:
Generación Dinámica: Se generan… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset_aeroespacial_cultural_completo.csv.CulturalPromptPortfolio
CulturalPromptPortfolio
tags: cultural, prompt, categorization
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'CulturalPromptPortfolio' dataset is curated to contain diverse image prompts that depict various ethnicities, aiming to serve as a valuable resource for machine learning practitioners focusing on cultural recognition and ethnically diverse representation. Each entry in the dataset is a textual prompt that is meant to… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/CulturalPromptPortfolio.cultural_dataset_teluguCulturalDataset_single_col
