CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HabibaAbderrahim /Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset Description This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations. It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.imagetranslationn<1K0 likes2.4k downloads1y agoHugging Face02aisingapore /Cultural-Evaluation-Kalahigated Kalahi Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset contains a MCQ-compatible version of the Kalahi dataset that is used in SEA-HELM. Supported Tasks and Leaderboards Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore. Languages Tagalog (tl)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi.textmultiple-choicen<1K0 likes890 downloads9mo agoHugging Face03haryoaw /cultural-spyfall Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game This dataset contains the game histories from Multicultural Spyfall, a dynamic benchmarking framework designed to evaluate the multilingual and multicultural capabilities of Large Language Models (LLMs). The dataset is introduced in the paper: Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game. Dataset Summary Multicultural Spyfall uses… See the full description on the dataset page: https://huggingface.co/datasets/haryoaw/cultural-spyfall.texttext-generation1K<n<10K1 likes203 downloads8mo agoHugging Face04PleIAs /Ukrainian-CulturalHeritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain. Dataset summary The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources. Curation method The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.tabulartext-generation10K<n<100K4 likes151 downloads3y agoHugging Face05SINAI /ALIA-es-cultural-heritage-synthetic-instructions Dataset Introduction The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains: 748,480 instances 629,682,398 tokens 25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.texttext-generation100K<n<1M0 likes79 downloads4mo agoHugging Face06boczkakaroly /trilingual-cultural-bias-redteaming-benchmark Trilingual Cultural Bias Red-Teaming Benchmark (HR–SR–HU) Overview This is a small qualitative benchmark for red-teaming large language models in Croatian (HR), Serbian (SR), and Hungarian (HU). The benchmark tests how models respond to provocative, culturally and historically loaded questions, when they are asked to role-play a patriotic citizen of a given country and answer in their own native language. The goal is not factual QA accuracy, but to observe reasoning… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/trilingual-cultural-bias-redteaming-benchmark.texttext-generationn<1K0 likes70 downloads9mo agoHugging Face07SINAI /ALIA-es-cultural-heritage-pairs Dataset Introduction The ALIA Spanish Cultural and Heritage Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen-style prompting workflow integrated in the ALIA encoders pipeline. It preserves provenance to the original document and passage while exposing controls such as question type and difficulty (ranging from high_school… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-pairs.textquestion-answering100K<n<1M0 likes70 downloads3mo agoHugging Face08Tushe /hausa-stem-reasoning-with-cultural-context Hausa STEM Reasoning with Cultural Context Abstract We present the first large-scale bilingual Hausa-English STEM reasoning dataset with deep cultural adaptation, containing 2,640 high-quality question-answer pairs translated from the STEM-Reasoning-Complex dataset. Our work introduces the "Shehin Malamin Kimiyya" (The Wise Scholar of Science) translation framework, which transforms Western scientific concepts into culturally-embedded Hausa explanations using systematic… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/hausa-stem-reasoning-with-cultural-context.textquestion-answering1K<n<10K1 likes69 downloads7mo agoHugging Face09AngelGabrielTroncoso /dataset-aeroespacial-cultural-somosnlp LATAM Aerospace History QA Descripción General LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica. El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos. La colección está especializada en: historia aeroespacial, programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.textquestion-answering100K<n<1M0 likes69 downloads4mo agoHugging Face10SINAI /ALIA-es-cultural-heritage Dataset Introduction The ALIA Spanish Cultural and Heritage Corpus is a strategic open data infrastructure designed to support research and innovation in digital humanities, cultural analytics, and Spanish-language NLP. It consolidates heterogeneous official and academic repositories into a single curated dataset, enabling broad and structured access to cultural heritage documentation from Spain. With 236,314 instances, 939,315,404 tokens and 100 source datasets, it provides a… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage.texttext-generation100K<n<1M0 likes68 downloads3mo agoHugging Face11SINAI /ALIA-es-cultural-heritage-triplets Dataset Introduction The dataset ALIA Spanish Cultural and Heritage Hard Negatives Corpus contains hard negatives for dense retrieval training generated from <query, passage> pairs contained in SINAI/ALIA-es-cultural-heritage-pairs.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish cultural heritage language. Hard negatives are passages that are semantically similar to a query but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-triplets.texttext-generation1M<n<10M0 likes54 downloads4mo agoHugging Face12abdullah693 /adaption-urdu-edu-cultural-reasoning This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-urdu_edu_cultural_reasoning This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.texttext-generation10K<n<100K0 likes51 downloads3mo agoHugging Face13saaga /yoruba-cultural-reasoning-blindspots# Yoruba Cultural Reasoning Blind Spots in Frontier Models ## Overview This dataset captures the "blind spots" of frontier base models when evaluating non-Western abstract reasoning, specifically focusing on Yoruba proverbs. It contains 10 diverse examples demonstrating "generative collapse" and "Cultural Hallucination." ## 1. Loading the Model This evaluation was conducted using a standard Google Colab environment with a T4 GPU. The model evaluated was `Qwen/Qwen2.5-1.5B`. It was loaded in… See the full description on the dataset page: https://huggingface.co/datasets/saaga/yoruba-cultural-reasoning-blindspots.texttext-generationn<1K1 likes33 downloads7mo agoHugging Face14StrataSynth /stratasynth-cross-cultural-negotiation StrataSynth Cross-Cultural Negotiation Benchmark Part of the StrataSynth Synthetic Identity Engineering corpus. 2,344 turns · 100 conversations · 50 GB + 50 US, pairwise matched · 24 columns per turn A controlled experiment, not just a corpus. The same two synthetic identities — a 47-year-old female VP of Procurement (buyer) and a 36-year-old male SaaS startup founder (vendor) — negotiate the same enterprise procurement deal 50 times in Great Britain and 50 times in the United… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-cross-cultural-negotiation.tabulartext-generation1K<n<10K0 likes29 downloads3d agoHugging Face15aliFurkan123 /cultural-questions-dataset Turkish General Knowledge & Trivia CoT Dataset (TR-GenK-CoT) TR-GenK-CoT is a high-quality, synthetic, and carefully curated Turkish dataset designed for Instruction Tuning and Chain of Thought (CoT) reasoning. It contains exactly 500 completely unique, non-repetitive general knowledge and trivia conversations with rich step-by-step thinking processes. The dataset is formatted using standard Chat Template formats (matching OpenAI/Hugging Face chat schemas) making it directly… See the full description on the dataset page: https://huggingface.co/datasets/aliFurkan123/cultural-questions-dataset.texttext-generationn<1K0 likes28 downloads2mo agoHugging Face16juice-cultural-eval /JuICE JuICE Sources Repository: https://anonymous.4open.science/r/JuICE HuggingFace: juice-cultural-eval/JuiCE About We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh), in both… See the full description on the dataset page: https://huggingface.co/datasets/juice-cultural-eval/JuICE.texttext-generation10K<n<100K1 likes25 downloads5mo agoHugging Face17Umkho-AI /SA_Cultural_Tribal_Practices SA Tribal & Cultural Practices Dataset Author: Minah Mojela (@minahmojela), Umkho-AI Dataset Summary This dataset contains 127 structured records documenting the cultural practices, customs, and identity histories of South Africa's major ethnic and population groups. It is a companion release to the South African History Dataset, built for the same reason: most AI models describe South African cultural practices using surface-level, externally-authored sources… See the full description on the dataset page: https://huggingface.co/datasets/Umkho-AI/SA_Cultural_Tribal_Practices.texttext-generationn<1K0 likes23 downloads2mo agoHugging Face18yale-cultural-heritage /synth-creation-date synth-creation-date Dataset Description This is a synthetic dataset for training date extraction models on cultural heritage and museum object descriptions. The dataset contains 792 samples of text descriptions paired with structured date information. Dataset Structure Data Fields prompt: Input text containing date information (string) completion: JSON string containing extracted dates with the following structure:{ "dates": [ { "type":… See the full description on the dataset page: https://huggingface.co/datasets/yale-cultural-heritage/synth-creation-date.texttext-generationn<1K0 likes18 downloads1y agoHugging Face19BuzzBlitz360A /Ukrainian-CulturalHeritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain. Dataset summary The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources. Curation method The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/BuzzBlitz360A/Ukrainian-CulturalHeritage-Books.tabulartext-generation10K<n<100K0 likes18 downloads8mo agoHugging Face20somosnlp-hackathon-2026 /patriae-cuban-cultural-appropriateness-prompts Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana como parte de su participación en el reto #HackathonSomosNLP 2026: Preferencias. Descripción General Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuban-cultural-appropriateness-prompts.texttext-classification1K<n<10K0 likes17 downloads4mo agoHugging Face21Patriae /patriae-cuban-cultural-appropriateness-prompts Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana Descripción General Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad cultural a través de las provincias de Cuba, permitiendo a investigadores y profesionales desarrollar y evaluar modelos que comprendan y respeten los matices culturales cubanos. Estadísticas del… See the full description on the dataset page: https://huggingface.co/datasets/Patriae/patriae-cuban-cultural-appropriateness-prompts.texttext-classification1K<n<10K0 likes16 downloads4mo agoHugging Face22Mohd12312 /darshana-cultural-heritage-dataset ???? DarShana India Cultural Heritage & Tourism Dataset This dataset powers the DarShana Living Cultural Traveler AI Platform, providing structured instruction-tuning examples and knowledge graph nodes for over 500+ Indian heritage destinations, seasonal fairs, living artisan clusters, and local cuisines. ?? Dataset Structure Each row in darshana_cultural_dataset.jsonl contains: instruction: The prompt task (e.g. Generate authentic cultural itinerary and seasonal… See the full description on the dataset page: https://huggingface.co/datasets/Mohd12312/darshana-cultural-heritage-dataset.texttext-generationn<1K0 likes11 downloads1mo agoHugging Face23Raniahossam33 /Arabic_cultural_dataset_with_openendedgated Arabic Cultural Dataset with MCQ and Open-Ended Answers Dataset Description This dataset contains culturally-aware questions in various Arabic dialects with BOTH multiple-choice options AND open-ended generation answers. It's designed to evaluate language models' understanding of cultural nuances across different Arabic-speaking regions in both structured (MCQ) and generative formats. Dataset Summary The dataset includes questions written in four major Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/Arabic_cultural_dataset_with_openended.textquestion-answering10K<n<100K0 likes10 downloads1y agoHugging Face24houssamboukhalfa /culturally_aligned_arabic_stories_subset_a 📚 Culturally Aligned Arabic Stories Dataset (Subset A) A curated 110-example subset of the Crafting Culturally Aligned Narratives dataset, designed for the development and evaluation of Arabic children’s story generation models aligned with Islamic and cultural values. ✨ Overview Language: Modern Standard Arabic (MSA) Samples: 110 prompt–response pairs Format: JSONL (id, language, prompt, response, source, license) Moral domains: honesty, courage, generosity… See the full description on the dataset page: https://huggingface.co/datasets/houssamboukhalfa/culturally_aligned_arabic_stories_subset_a.texttext-generationn<1K0 likes8 downloads11mo agoHugging Face25zox-BT /gemma-2b-cameroon-cultural-blindspots Gemma-2b Cameroon Cultural Blindspots This dataset highlights the "blind spots" of the Google Gemma-2-2b base model regarding Cameroonian culture, geography, and local languages. 1. Model Tested Model Name: google/gemma-2-2b Type: Base Model (Pre-trained) 2. Loading Procedure The model was loaded using the transformers library on a Google Colab T4 GPU: from transformers import AutoTokenizer, AutoModelForCausalLM import torch model_id = "google/gemma-2-2b"… See the full description on the dataset page: https://huggingface.co/datasets/zox-BT/gemma-2b-cameroon-cultural-blindspots.texttext-generationn<1K0 likes6 downloads7mo agoHugging Face26palestinian-kg /palestinian-cultural-knowledgegated Palestinian Cultural Knowledge Corpus v0.2.0 — supersedes the earlier data/wikipedia_ar/ v0.1.0 partial upload (484 Arabic Wikipedia documents only). This release expands to the full 5-source corpus below and moves the data to data/full_corpus/. A multi-source Arabic/English text corpus about Palestinian history, culture, and heritage, built for the Palestinian Cultural Knowledge Platform — a RAG + knowledge-graph research project. 882 documents, ~890K words, collected and… See the full description on the dataset page: https://huggingface.co/datasets/palestinian-kg/palestinian-cultural-knowledge.tabulartext-classificationn<1K1 likes6 downloads2mo agoHugging Face27farabi-lab /Cultural-Quizzesgated 🇰🇿 Kazakh Cultural Inquiry and Question Generation 📖 Overview This dataset contains 500 samples focused on the ability to generate structured, relevant, and inquisitive content based on Kazakh cultural topics. The primary task demonstrated in this dataset is Question Generation (QG), where the model takes a broad topic or user interest and produces a series of detailed, numbered questions to guide further research or discussion. 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Cultural-Quizzes.texttext-generationn<1K0 likes5 downloads2mo agoHugging Face28AngelGabrielTroncoso /dataset_aeroespacial_cultural_completo.csv 🚀 LATAM Aerospace Cultural QA Dataset culturalmente alineado para modelos conversacionales en español y portugués, especializado en historia aeroespacial iberoamericana. Desarrollado para el #HackathonSomosNLP 2026 — ¿Son los LLMs realmente multiculturales? 🛠 Metodología y Pipeline de Construcción La versión actual del dataset ha sido refinada mediante un pipeline automatizado diseñado para maximizar la calidad y la diversidad cultural: Generación Dinámica: Se generan… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset_aeroespacial_cultural_completo.csv.texttext-generation1K<n<10K0 likes4 downloads4mo agoHugging Face29borvo-sc /patrimonio-cultural-PTgatedtexttext-generationn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.