datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AraDiCE-Culture
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository we show the cultural split of the data
Evaluation
We have used lm-harness eval framework to for the… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE-Culture.kencorpus_sw_culture
KenCorpus Swahili Culture Subset
A filtered subset of Kencorpus/KenCorpus_audio,
containing only rows where language=Swahili and genre=Culture (37 clips).
Audio files are in audio/, indexed by kencorpus_sw_culture.jsonl with path and duration fields,
following the layout of kyutai/DailyTalkContiguous.
ai-culture-multilingual-json-dolma
AI-Culture Multilingual JSON + DOLMA Corpus
16M words · 12 languages · CC-BY-4.0
The AI-Culture corpus contains 5K articles providing comprehensive philosophical and cultural content, exploring the intersection of technology, artificial intelligence, and human culture, perfectly aligned across 12 languages. All content maintains identical parallel structure across translations with zero duplication and editor-curated quality.
This project is maintained by a non-profit digital… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/ai-culture-multilingual-json-dolma.multilingual-idioms-for-cultureculture-mt-repro-artifacts
CULTURE-MT Reproduction Artifacts — ICML 2026 #33401
Reproduction artifacts for "Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC" (OpenReview 7xppFNbcXM, arXiv 2605.25626), produced for the ICML-2026 Agent Repro challenge.
📓 Logbook (full write-up, claim by claim): https://huggingface.co/spaces/Manaiakalani/culture-mt-repro
Contents
Path
What
data/CULTURE-MT_test.jsonl
1,002 source Chinese UGC notes (from… See the full description on the dataset page: https://huggingface.co/datasets/Manaiakalani/culture-mt-repro-artifacts.PT-Culture_Data
Portuguese Cultural SFT Dataset
A supervised fine-tuning dataset for European Portuguese (PT-PT) cultural knowledge, built to teach models the traditions, figures, places, and expressions of Portuguese culture.
The dataset contains 216,832 examples across two versions in conversational SFT format, organised into ten cultural domains:
Domain
v1 (95,815)
v2 (121,017)
Total
Personalities
36,370
62,788
99,158
Audiovisual
15,786
26,857
42,643
Heritage
8,020… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/PT-Culture_Data.CULTURE-MT
CULTURE-MT: Beyond Literal Translation — Evaluating Cultural Effectiveness in Social Media UGC
CULTURE-MT is a benchmark for evaluating CULtural Transmission and UGC-specific emotion REsonance in Chinese-to-English social media translation. It consists of 1,002 user-generated notes (UGC) spanning 14 content domains, presented at ICML 2026.
📄 Paper: Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC
🌟 Why CULTURE-MT?
Standard machine… See the full description on the dataset page: https://huggingface.co/datasets/Wulinjuan/CULTURE-MT.adaption-african-culture-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-african_culture_qa
This dataset consists of question-and-answer pairs focusing on African cultural traditions, festivals, religious practices, and musical instruments. The content covers diverse regions and ethnic groups, including the Yoruba, Twi, Luganda, Gishu, and Nyankore peoples. Entries vary in format from true/false statements to multiple-choice questions and descriptive… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-african-culture-qa.kaz-culture-qa-15k
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
kaz_culture_qa_15k
This dataset contains question-answer pairs in the Kazakh language focused on Kazakh culture, traditions, literature, and customs. The samples cover topics such as Mukhtar Auezov's works, traditional jewelry, rituals like Nauryz blessings, and social norms. It is structured for question-answering tasks with approximately 16,000 examples split into training… See the full description on the dataset page: https://huggingface.co/datasets/shayekh/kaz-culture-qa-15k.adaption-japanese-instruction-culture
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-japanese_instruction_culture
This dataset contains Japanese instruction-tuning samples covering summarization, business communication, math, coding, and creative writing. A significant portion is dedicated to deep cultural explanations, including concepts like wabi-sabi, amae, and hare-ke, as well as detailed guidance on keigo honorifics. The content reflects high-context Japanese… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-japanese-instruction-culture.football-culture-reasoning
Football Culture Reasoning Bench v1
Expert-graded evaluation of LLM reasoning about football (soccer) fandom culture: subcultural concepts, non-Western specificity (Japan / J.League, South America, Asia), macro-sociological context, and stereotype avoidance.
This is a small, deliberately hard proof set (20 items). The failures it probes are cultural, not linguistic — candidate answers are fluent and confident, but stale, West-centric, or normatively preachy.
Task… See the full description on the dataset page: https://huggingface.co/datasets/Kablue/football-culture-reasoning.turkish-culture-dataK-Culture-Desc
K-Culture Contextual Understanding Benchmark
A Korean cultural understanding benchmark dataset featuring 530 scenario-based multiple-choice questions designed to evaluate models' contextual understanding of Korean culture.
Dataset Description
K-Culture Contextual Understanding Benchmark is a dataset designed to evaluate LLMs' understanding of Korean cultural contexts through realistic scenarios and dialogues. Each item contains a cultural description, a scenario depicting… See the full description on the dataset page: https://huggingface.co/datasets/SOGANG-ISDS/K-Culture-Desc.filipino_morals_culture_elementary_periodical_examsPT-Culture_DataCultureNuc_ftvi-culture-composite-annotatedadaption-african-culture-and-history
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-african_culture_and_history
This dataset consists of instruction and response pairs covering African history, geography, and societal customs. It explores the origins, practices, and geographic locations of diverse African kingdoms and cultures. Content highlights oral traditions, cultural identity, and human geography across various regions of the continent.
Dataset size… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-african-culture-and-history.Konoturkish-culture-data-2EthicsAI-K-Culture-Desc
한국 문화 맥락 이해 벤치마크 (K-Culture Contextual Understanding Benchmark)
한국 문화에 대한 모델의 맥락적 이해도를 평가하기 위해 설계된 530개의 시나리오 기반 객관식 질문을 포함하는 한국 문화 이해 벤치마크 데이터셋입니다.
데이터셋 설명
한국 문화 맥락 이해 벤치마크는 실생활 시나리오와 대화를 통해 거대언어모델(LLM)의 한국 문화 맥락 이해 능력을 평가하기 위해 설계된 데이터셋입니다. 각 항목은 문화적 설명, 실제 상황을 묘사한 시나리오, 그리고 상세한 해설이 포함된 객관식 질문으로 구성되어 있습니다.
언어: 한국어 (ko)
크기: 530개 항목
작업: 객관식 질의응답 (MCQA)
버전: v1.0
데이터 구조
데이터 필드
필드명
타입
설명
DescriptionID
int
각 항목의 고유 식별자
Description
string
한국… See the full description on the dataset page: https://huggingface.co/datasets/saltlux/EthicsAI-K-Culture-Desc.Bengal_CultureConnor-Data-AEX-art_culturemix-culture-food-train
Mix Culture Food Dataset — Train Split
This repository contains the training split for the mix-culture food fine-tuning project. It includes:
train_dataset_hf.jsonl: 5,768 multimodal training samples.
images/: Only the images referenced in the JSONL, organized by their source dataset.
Each JSONL entry follows the Swift "messages + images" schema. Questions mention either the food itself or the left food (for multi-dish compositions). Answers are deterministic lookups from… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/mix-culture-food-train.mix-culture-food-test
Mix Culture Food Dataset — Test Split
Evaluation-only split corresponding to the mix-culture food fine-tuning corpus.
test_dataset_hf.jsonl: 7,092 validation samples.
images/: Only the referenced evaluation images, organized by source dataset.
Structure mirrors the train split; each line in the JSONL contains messages, images, and metadata fields compatible with Swift's AutoPreprocessor.
Directory Layout
.
├── README.md
├── test_dataset_hf.jsonl
└── images/
├──… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/mix-culture-food-test.Connor-Data-Art_culturegreek_culture_rag_datasetequiDesmongolian-cultureequiLM
