datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
viet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.culturalmoment-benchmarkPaper | Project Page | Leaderboard | Walkthrough | SCB, the image predecessor
Cultural Moment Benchmark (CMB)
Evaluating Video Cultural Reasoning and Grounding in Southeast Asia
CMB evaluates how vision-language models reason about cultural moments in video across Southeast Asia. Each concept is tested in three stages: naming the concept, recognizing it visually in video, and temporally localizing its sub-events, under three context modes (Reset, Carry… See the full description on the dataset page: https://huggingface.co/datasets/Multimedia-SMU/culturalmoment-benchmark.ACVA-Arabic-Cultural-Value-Alignment
About ArabicCulture
The ArabicCulture dataset was generated by gpt3.5 and contains 8000+ True and False questions.The dataset contains questions from 58 different areas.In the answers, "True" accounted for 59.62%, and "False" accounted for 40.38%
data-all
It contains 8000+ data, and we took 5 data from each area as few-shot data.
data-select
We asked two Arabs to judge 4000 of all the data for us, and we left data that two Arabs both thought were good. Finally… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ACVA-Arabic-Cultural-Value-Alignment.ALIA-es-cultural-heritage-synthetic-instructions
Dataset Introduction
The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision.
It contains:
748,480 instances
629,682,398 tokens
25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.ALIA-es-cultural-heritage-pairs
Dataset Introduction
The ALIA Spanish Cultural and Heritage Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen-style prompting workflow integrated in the ALIA encoders pipeline.
It preserves provenance to the original document and passage while exposing controls such as question type and difficulty (ranging from high_school… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-pairs.dataset-aeroespacial-cultural-somosnlp
LATAM Aerospace History QA
Descripción General
LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica.
El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos.
La colección está especializada en:
historia aeroespacial,
programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.shadow-puppet-outputadaption-urdu-edu-cultural-reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-urdu_edu_cultural_reasoning
This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.CulturalDrive-Bench-v2
CulturalDrive-Bench
An open-ended VQA benchmark that tests whether a vision-language model can infer
the driving region from a dashcam frame alone — no country label is given —
and then apply that region's traffic rules.
40,169 items across 8 countries. Every item carries the specific traffic
rule its answer depends on.
Frames are bundled under images/, one tar per source dataset, downscaled
to a 1280px long edge. 49,765 images, 9.5 GB. The originals stay with their
upstream… See the full description on the dataset page: https://huggingface.co/datasets/gray311/CulturalDrive-Bench-v2.gsm8k-indic-cultural
GSM8K Indic Cultural Adaptation
Dataset Summary
GSM8K Indic Cultural Adaptation is a culturally localized version of the GSM8K test split, designed to evaluate the robustness of mathematical reasoning models under culturally adapted problem formulations.
The dataset preserves the underlying mathematical reasoning of the original GSM8K benchmark while adapting questions to an Indian context. Depending on the variant, this includes replacing culturally specific… See the full description on the dataset page: https://huggingface.co/datasets/kiranpradeep/gsm8k-indic-cultural.indian-cultural-datasetcultural-questions-dataset
Turkish General Knowledge & Trivia CoT Dataset (TR-GenK-CoT)
TR-GenK-CoT is a high-quality, synthetic, and carefully curated Turkish dataset designed for Instruction Tuning and Chain of Thought (CoT) reasoning. It contains exactly 500 completely unique, non-repetitive general knowledge and trivia conversations with rich step-by-step thinking processes.
The dataset is formatted using standard Chat Template formats (matching OpenAI/Hugging Face chat schemas) making it directly… See the full description on the dataset page: https://huggingface.co/datasets/aliFurkan123/cultural-questions-dataset.GSM8K-cultural
GSM8K_Cultural
This dataset is part of our investigation into how cultural context influences the performance of large language models (LLMs) on mathematical problems.
The methodology for creating this dataset is detailed in the research paper:Lost in Cultural Translation: Do LLMs Struggle with Math Across
Cultural Contexts?. Code can also be found on Github
What is this dataset
This dataset consists of six cultural variants of the GSM8K test set. Each variant retains… See the full description on the dataset page: https://huggingface.co/datasets/abedk/GSM8K-cultural.SA_Cultural_Tribal_Practices
SA Tribal & Cultural Practices Dataset
Author: Minah Mojela (@minahmojela), Umkho-AI
Dataset Summary
This dataset contains 127 structured records documenting the cultural practices,
customs, and identity histories of South Africa's major ethnic and population groups.
It is a companion release to the South African History Dataset,
built for the same reason: most AI models describe South African cultural practices
using surface-level, externally-authored sources… See the full description on the dataset page: https://huggingface.co/datasets/Umkho-AI/SA_Cultural_Tribal_Practices.adaption-african-cultural-qa
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
african_cultural_qa
This dataset contains 1,000 question-and-answer pairs exploring diverse aspects of African culture, including mythology, language, traditional leadership, and social philosophies like Ubuntu. Each entry features a prompt, a detailed completion, and associated metadata fields for reasoning and topic classification. The content focuses on the historical… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/adaption-african-cultural-qa.en-si-translation-cultural-idioms-500
En Si Translation Cultural Idioms 500
Dataset Summary
English-Sinhala Idioms and Cultural Expressions dataset featuring ~500 items designed to teach semantic context mapping over literal word-for-word translation transitions.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 500
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-cultural-idioms-500.Cultural-Safety-Telugu
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-telugu_safety_prompts
A dataset of Telugu-language prompts labeled as SAFE or UNSAFE based on content safety. It includes text prompts and their corresponding safety annotations, intended for training or evaluating content moderation systems. The dataset focuses on identifying harmful or inappropriate content in Telugu.
Dataset size
There are 13,920… See the full description on the dataset page: https://huggingface.co/datasets/salmankhanpm/Cultural-Safety-Telugu.darshana-cultural-heritage-dataset
???? DarShana India Cultural Heritage & Tourism Dataset
This dataset powers the DarShana Living Cultural Traveler AI Platform, providing structured instruction-tuning examples and knowledge graph nodes for over 500+ Indian heritage destinations, seasonal fairs, living artisan clusters, and local cuisines.
?? Dataset Structure
Each row in darshana_cultural_dataset.jsonl contains:
instruction: The prompt task (e.g. Generate authentic cultural itinerary and seasonal… See the full description on the dataset page: https://huggingface.co/datasets/Mohd12312/darshana-cultural-heritage-dataset.cultural_awareness_mcqculturally_aligned_arabic_stories_subset_a
📚 Culturally Aligned Arabic Stories Dataset (Subset A)
A curated 110-example subset of the Crafting Culturally Aligned Narratives dataset, designed for the development and evaluation of Arabic children’s story generation models aligned with Islamic and cultural values.
✨ Overview
Language: Modern Standard Arabic (MSA)
Samples: 110 prompt–response pairs
Format: JSONL (id, language, prompt, response, source, license)
Moral domains: honesty, courage, generosity… See the full description on the dataset page: https://huggingface.co/datasets/houssamboukhalfa/culturally_aligned_arabic_stories_subset_a.vi-cultural-benchmarkCross-Cultural-Multilingual-Teachingagentic-lux-personkerala-cultural-kg-seedCulturalLLMs-DPONotes: Data extracted from World Value Survey wave 7th.
sft: SFT data
DPO: Paired data for all permutations of options
DPO-refined_input: The paired data of all the options are combined in pairs. The prompts are modified to require the model to choose between the two paired options.
DPO-refined_cr: The pairing data of the options are combined in pairs. The pairing method is modified to ensure that all chosen options are the options with the highest probability, and rejected options are all… See the full description on the dataset page: https://huggingface.co/datasets/alec-x/CulturalLLMs-DPO.cultural-rdfpatrimonio-cultural-PT
