CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IAmFuch /viet-cultural-vqa 🇻🇳 Vietnamese Cultural VQA Dataset 📖 Dataset Description The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering. 🎯 Dataset Summary 📊 Total Images: 28,505 high-quality cultural images 💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.imagevisual-question-answering10K<n<100K0 likes534 downloads5mo agoHugging Face02Multimedia-SMU /culturalmoment-benchmarkPaper | Project Page | Leaderboard | Walkthrough | SCB, the image predecessor Cultural Moment Benchmark (CMB) Evaluating Video Cultural Reasoning and Grounding in Southeast Asia CMB evaluates how vision-language models reason about cultural moments in video across Southeast Asia. Each concept is tested in three stages: naming the concept, recognizing it visually in video, and temporally localizing its sub-events, under three context modes (Reset, Carry… See the full description on the dataset page: https://huggingface.co/datasets/Multimedia-SMU/culturalmoment-benchmark.textvisual-question-answeringn<1K1 likes187 downloads10d agoHugging Face03FreedomIntelligence /ACVA-Arabic-Cultural-Value-Alignment About ArabicCulture The ArabicCulture dataset was generated by gpt3.5 and contains 8000+ True and False questions.The dataset contains questions from 58 different areas.In the answers, "True" accounted for 59.62%, and "False" accounted for 40.38% data-all It contains 8000+ data, and we took 5 data from each area as few-shot data. data-select We asked two Arabs to judge 4000 of all the data for us, and we left data that two Arabs both thought were good. Finally… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ACVA-Arabic-Cultural-Value-Alignment.text1K<n<10K8 likes108 downloads3y agoHugging Face04SINAI /ALIA-es-cultural-heritage-synthetic-instructions Dataset Introduction The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains: 748,480 instances 629,682,398 tokens 25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.texttext-generation100K<n<1M0 likes79 downloads4mo agoHugging Face05SINAI /ALIA-es-cultural-heritage-pairs Dataset Introduction The ALIA Spanish Cultural and Heritage Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen-style prompting workflow integrated in the ALIA encoders pipeline. It preserves provenance to the original document and passage while exposing controls such as question type and difficulty (ranging from high_school… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-pairs.textquestion-answering100K<n<1M0 likes70 downloads3mo agoHugging Face06AngelGabrielTroncoso /dataset-aeroespacial-cultural-somosnlp LATAM Aerospace History QA Descripción General LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica. El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos. La colección está especializada en: historia aeroespacial, programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.textquestion-answering100K<n<1M0 likes69 downloads4mo agoHugging Face07yale-cultural-heritage /shadow-puppet-outputtext1K<n<10K0 likes61 downloads1y agoHugging Face08abdullah693 /adaption-urdu-edu-cultural-reasoning This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-urdu_edu_cultural_reasoning This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.texttext-generation10K<n<100K0 likes51 downloads3mo agoHugging Face09gray311 /CulturalDrive-Bench-v2 CulturalDrive-Bench An open-ended VQA benchmark that tests whether a vision-language model can infer the driving region from a dashcam frame alone — no country label is given — and then apply that region's traffic rules. 40,169 items across 8 countries. Every item carries the specific traffic rule its answer depends on. Frames are bundled under images/, one tar per source dataset, downscaled to a 1280px long edge. 49,765 images, 9.5 GB. The originals stay with their upstream… See the full description on the dataset page: https://huggingface.co/datasets/gray311/CulturalDrive-Bench-v2.documentvisual-question-answering10K<n<100K0 likes46 downloads10d agoHugging Face10kiranpradeep /gsm8k-indic-cultural GSM8K Indic Cultural Adaptation Dataset Summary GSM8K Indic Cultural Adaptation is a culturally localized version of the GSM8K test split, designed to evaluate the robustness of mathematical reasoning models under culturally adapted problem formulations. The dataset preserves the underlying mathematical reasoning of the original GSM8K benchmark while adapting questions to an Indian context. Depending on the variant, this includes replacing culturally specific… See the full description on the dataset page: https://huggingface.co/datasets/kiranpradeep/gsm8k-indic-cultural.tabularquestion-answering1K<n<10K0 likes36 downloads2mo agoHugging Face11ombhojane /indian-cultural-datasettextn<1K0 likes28 downloads2y agoHugging Face12aliFurkan123 /cultural-questions-dataset Turkish General Knowledge & Trivia CoT Dataset (TR-GenK-CoT) TR-GenK-CoT is a high-quality, synthetic, and carefully curated Turkish dataset designed for Instruction Tuning and Chain of Thought (CoT) reasoning. It contains exactly 500 completely unique, non-repetitive general knowledge and trivia conversations with rich step-by-step thinking processes. The dataset is formatted using standard Chat Template formats (matching OpenAI/Hugging Face chat schemas) making it directly… See the full description on the dataset page: https://huggingface.co/datasets/aliFurkan123/cultural-questions-dataset.texttext-generationn<1K0 likes28 downloads2mo agoHugging Face13abedk /GSM8K-cultural GSM8K_Cultural This dataset is part of our investigation into how cultural context influences the performance of large language models (LLMs) on mathematical problems. The methodology for creating this dataset is detailed in the research paper:Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?. Code can also be found on Github What is this dataset This dataset consists of six cultural variants of the GSM8K test set. Each variant retains… See the full description on the dataset page: https://huggingface.co/datasets/abedk/GSM8K-cultural.text1K<n<10K2 likes23 downloads1y agoHugging Face14Umkho-AI /SA_Cultural_Tribal_Practices SA Tribal & Cultural Practices Dataset Author: Minah Mojela (@minahmojela), Umkho-AI Dataset Summary This dataset contains 127 structured records documenting the cultural practices, customs, and identity histories of South Africa's major ethnic and population groups. It is a companion release to the South African History Dataset, built for the same reason: most AI models describe South African cultural practices using surface-level, externally-authored sources… See the full description on the dataset page: https://huggingface.co/datasets/Umkho-AI/SA_Cultural_Tribal_Practices.texttext-generationn<1K0 likes23 downloads2mo agoHugging Face15Svngoku /adaption-african-cultural-qa This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. african_cultural_qa This dataset contains 1,000 question-and-answer pairs exploring diverse aspects of African culture, including mythology, language, traditional leadership, and social philosophies like Ubuntu. Each entry features a prompt, a detailed completion, and associated metadata fields for reasoning and topic classification. The content focuses on the historical… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/adaption-african-cultural-qa.textn<1K0 likes22 downloads5mo agoHugging Face16SAWithanage /en-si-translation-cultural-idioms-500 En Si Translation Cultural Idioms 500 Dataset Summary English-Sinhala Idioms and Cultural Expressions dataset featuring ~500 items designed to teach semantic context mapping over literal word-for-word translation transitions. Engineering Pipeline Parameters Language Pair: English (en) to Sinhala (si) Total Valid Token Rows: 500 Internal Storage Structure: Single-File data.json Upstream Source Attribution This specific sub-split was compiled and… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-cultural-idioms-500.texttranslationn<1K0 likes16 downloads4mo agoHugging Face17salmankhanpm /Cultural-Safety-Telugu This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-telugu_safety_prompts A dataset of Telugu-language prompts labeled as SAFE or UNSAFE based on content safety. It includes text prompts and their corresponding safety annotations, intended for training or evaluating content moderation systems. The dataset focuses on identifying harmful or inappropriate content in Telugu. Dataset size There are 13,920… See the full description on the dataset page: https://huggingface.co/datasets/salmankhanpm/Cultural-Safety-Telugu.text10K<n<100K0 likes14 downloads3mo agoHugging Face18Mohd12312 /darshana-cultural-heritage-dataset ???? DarShana India Cultural Heritage & Tourism Dataset This dataset powers the DarShana Living Cultural Traveler AI Platform, providing structured instruction-tuning examples and knowledge graph nodes for over 500+ Indian heritage destinations, seasonal fairs, living artisan clusters, and local cuisines. ?? Dataset Structure Each row in darshana_cultural_dataset.jsonl contains: instruction: The prompt task (e.g. Generate authentic cultural itinerary and seasonal… See the full description on the dataset page: https://huggingface.co/datasets/Mohd12312/darshana-cultural-heritage-dataset.texttext-generationn<1K0 likes11 downloads1mo agoHugging Face19potsawee /cultural_awareness_mcqtext1K<n<10K0 likes10 downloads2y agoHugging Face20houssamboukhalfa /culturally_aligned_arabic_stories_subset_a 📚 Culturally Aligned Arabic Stories Dataset (Subset A) A curated 110-example subset of the Crafting Culturally Aligned Narratives dataset, designed for the development and evaluation of Arabic children’s story generation models aligned with Islamic and cultural values. ✨ Overview Language: Modern Standard Arabic (MSA) Samples: 110 prompt–response pairs Format: JSONL (id, language, prompt, response, source, license) Moral domains: honesty, courage, generosity… See the full description on the dataset page: https://huggingface.co/datasets/houssamboukhalfa/culturally_aligned_arabic_stories_subset_a.texttext-generationn<1K0 likes8 downloads11mo agoHugging Face21nygdon /vi-cultural-benchmarktextmultiple-choice1K<n<10K0 likes8 downloads5mo agoHugging Face22Quay1k /Cross-Cultural-Multilingual-Teachingtextn<1K0 likes7 downloads3y agoHugging Face23yale-cultural-heritage /agentic-lux-persontextn<1K0 likes7 downloads5mo agoHugging Face24kaumudi-ai /kerala-cultural-kg-seedtextn<1K0 likes4 downloads3mo agoHugging Face25alec-x /CulturalLLMs-DPOgatedNotes: Data extracted from World Value Survey wave 7th. sft: SFT data DPO: Paired data for all permutations of options DPO-refined_input: The paired data of all the options are combined in pairs. The prompts are modified to require the model to choose between the two paired options. DPO-refined_cr: The pairing data of the options are combined in pairs. The pairing method is modified to ensure that all chosen options are the options with the highest probability, and rejected options are all… See the full description on the dataset page: https://huggingface.co/datasets/alec-x/CulturalLLMs-DPO.textquestion-answering100K<n<1M1 likes2 downloads1y agoHugging Face26duohub-ai /cultural-rdfgatedtextn<1K1 likes1 downloads2y agoHugging Face27borvo-sc /patrimonio-cultural-PTgatedtexttext-generationn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.