CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes30 downloads2y agoHugging Face02Rabe3 /sera-phase2-saudi-dialect-rag SERA Saudi Dialect RAG Dataset (Phase 2 - Domain) Domain-specific RAG fine-tuning dataset for Saudi Arabic dialect, focused on Saudi Electricity Regulatory Authority (SERA) documents. Format LlamaFactory Alpaca format: Field Description instruction System prompt + real document chunk as context + question in Saudi dialect input Always empty output Answer in Saudi dialect Usage with LlamaFactory Copy the JSON files into your LlamaFactory… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/sera-phase2-saudi-dialect-rag.textquestion-answering1K<n<10K0 likes29 downloads7mo agoHugging Face03Rabe3 /saudi-dialect-rag Saudi Dialect RAG Fine-Tuning Dataset A RAG-formatted fine-tuning dataset for Saudi Arabic dialect, built from HeshamHaroon/saudi-dialect-conversations. Format Each example follows the LlamaFactory Alpaca format: Field Description instruction System prompt + MSA context paragraph + optional conversation history + question input Always empty string output Assistant reply in Saudi dialect How it was built Loaded source multi-turn Saudi… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/saudi-dialect-rag.textquestion-answering10K<n<100K0 likes25 downloads7mo agoHugging Face04yusufbaykaloglu /Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT Turkish Dialectical Reasoning Dataset (Sokrates-ToT) The Turkish Dialectical Reasoning Dataset (Sokrates-ToT) is a collection structured in a Tree-of-Thought (ToT) format, based on a multi-persona and dialectical reasoning framework.Inspired by Socrates' method of dialogue, it facilitates deep analysis of complex and multidimensional issues by having AI personas with different expertise interact and ultimately reach a final synthesis. Purpose of the Dataset This… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT.textquestion-answering1K<n<10K0 likes17 downloads1y agoHugging Face05KIND-Dataset /Open-ended_Questions_dialectal_data Dataset Summary A collection of open-ended questions that was provided to the data marathon competitors to populate KIND dataset. It was designed to elicit longer responses cultural and context-rich sentences. For more details, please check the paper The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection Citation Information @inproceedings{yamani-etal-2024-kind, title = "The {KIND} Dataset: A Social Collaboration Approach for Nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/Open-ended_Questions_dialectal_data.textquestion-answeringn<1K0 likes16 downloads3y agoHugging Face06somosnlp-hackathon-2025 /exam_zh_multitopic_dialect_culture exam_zh_multitopic_dialect_culture This dataset contains 300 multiple-choice questions (MCQs) from a variety of Mandarin-based assessments, spanning both regional dialect comprehension and cultural/general knowledge. 📚 Description The questions come from publicly available Chinese-language exams and quizzes, and fall into two major categories: 🗣️ Regional Dialect Tests These assess language understanding across major Chinese dialects and topolects: Hakka… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/exam_zh_multitopic_dialect_culture.tabularmultiple-choicen<1K0 likes15 downloads1y agoHugging Face07saeedbark /hadrami-arabic-dialect-dataset Hadrami Arabic Dialect Dataset A structured dataset of 1,000 entries from the Hadrami Arabic dialect (spoken primarily in the Hadramawt region of Yemen). Each entry includes the dialectal word alongside its Modern Standard Arabic (MSA/Fusha) equivalent, linguistic metadata, usage examples, proverbs, and semantic tags. 🔊 Roadmap: Audio pronunciations for each entry are planned for a future release. Dataset Summary Field Value Entries 1,000 Language… See the full description on the dataset page: https://huggingface.co/datasets/saeedbark/hadrami-arabic-dialect-dataset.texttext-classification1K<n<10K1 likes13 downloads5mo agoHugging Face08yukebrillianth /east_java_dialect_instruct Complaints From The East Javanese Dialect community This dataset created manually by humans with reference to public complaints in the comments column of the local government's Instagram account and another platform like X and TikTok Comments. textquestion-answeringn<1K0 likes12 downloads1y agoHugging Face09KareemBb /Jordanian-Dialect-Instruct-QA Dataset Card: Jordanian-Dialect-Instruct-QA Overview This dataset contains 1500 academic Q&A pairs for Jordanian universities, written in the Jordanian Arabic dialect. It is used to fine-tune models to provide natural, localized responses to academic inquiries and more. This dataset was used to fine-tune Qwen2.5-7B-Instruct-Jordanian. You can visit the model page to download the GGUF weights and the custom Ollama Modelfile. Dataset Structure Each sample… See the full description on the dataset page: https://huggingface.co/datasets/KareemBb/Jordanian-Dialect-Instruct-QA.textquestion-answering1K<n<10K1 likes8 downloads4mo agoHugging Face10farabi-lab /Understanding_Dialect_Textgated 🇰🇿 Kazakh Dialect Analysis and Standardization Dataset Dataset Summary Kazakh Dialect Analysis and Standardization Dataset is a Kazakh-language linguistics dataset designed for dialect identification, dialectal feature analysis, and normalization into standard Kazakh. The dataset contains instruction-style prompts asking the model to identify dialectal or regional language features in a given Kazakh text and explain how they can be converted into standard… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Understanding_Dialect_Text.texttext-generationn<1K0 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.