datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.trilingual-cultural-bias-redteaming-benchmark
Trilingual Cultural Bias Red-Teaming Benchmark (HR–SR–HU)
Overview
This is a small qualitative benchmark for red-teaming large language models in Croatian (HR), Serbian (SR), and Hungarian (HU).
The benchmark tests how models respond to provocative, culturally and historically loaded questions, when they are asked to role-play a patriotic citizen of a given country and answer in their own native language.
The goal is not factual QA accuracy, but to observe reasoning… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/trilingual-cultural-bias-redteaming-benchmark.JuICE
JuICE
Sources
Repository: https://anonymous.4open.science/r/JuICE
HuggingFace: juice-cultural-eval/JuiCE
About
We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh), in both… See the full description on the dataset page: https://huggingface.co/datasets/juice-cultural-eval/JuICE.patriae-cuban-cultural-appropriateness-prompts
Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana como parte de su participación en el reto #HackathonSomosNLP 2026: Preferencias.
Descripción General
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuban-cultural-appropriateness-prompts.patriae-cuban-cultural-appropriateness-prompts
Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana
Descripción General
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad cultural a través de las provincias de Cuba, permitiendo a investigadores y profesionales desarrollar y evaluar modelos que comprendan y respeten los matices culturales cubanos.
Estadísticas del… See the full description on the dataset page: https://huggingface.co/datasets/Patriae/patriae-cuban-cultural-appropriateness-prompts.gemma-2b-cameroon-cultural-blindspots
Gemma-2b Cameroon Cultural Blindspots
This dataset highlights the "blind spots" of the Google Gemma-2-2b base model regarding Cameroonian culture, geography, and local languages.
1. Model Tested
Model Name: google/gemma-2-2b
Type: Base Model (Pre-trained)
2. Loading Procedure
The model was loaded using the transformers library on a Google Colab T4 GPU:
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "google/gemma-2-2b"… See the full description on the dataset page: https://huggingface.co/datasets/zox-BT/gemma-2b-cameroon-cultural-blindspots.dataset_aeroespacial_cultural_completo.csv
🚀 LATAM Aerospace Cultural QA
Dataset culturalmente alineado para modelos conversacionales en español y portugués, especializado en historia aeroespacial iberoamericana.
Desarrollado para el #HackathonSomosNLP 2026 — ¿Son los LLMs realmente multiculturales?
🛠 Metodología y Pipeline de Construcción
La versión actual del dataset ha sido refinada mediante un pipeline automatizado diseñado para maximizar la calidad y la diversidad cultural:
Generación Dinámica: Se generan… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset_aeroespacial_cultural_completo.csv.
