CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mesolitica /Malay-Dialect-Instructions Malay dialect instruction including coding Negeri Sembilan QA public transport QA, Coding CUDA coding, Kedah QA infra QA, Coding Rust coding, Kelantan QA Najib Razak QA, Coding Go coding, Perak QA Anwar Ibrahim QA, Coding SQL coding, Pahang QA Pendatang asing QA, Coding Typescript coding, Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.texttext-generation10K<n<100K5 likes1.8k downloads2y agoHugging Face02ISLAM-PO /arab-dialects-20-countries-3m Arab Dialects Dataset - 20 Countries A large-scale Arabic dialects dataset covering 20 Arab countries, 7 content types per country, 3,000,000 records, 140 JSONL files, 12.07 GB. UTF-8 JSONL, ready for Hugging Face Datasets. 1. Contents 1. Contents 2. Dataset Summary 3. Repository Map 4. Countries Table (20 folders) 5. Data Types Table (7 files) 6. Record Schema 7. Loading and Usage 8. Generation and Reproduction 9. Considerations and Limitations 10. Contributors… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.texttext-generation1M<n<10M0 likes585 downloads19d agoHugging Face03HeshamHaroon /saudi-dialect-conversations Saudi Najdi Dialect Conversations A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models. Dataset Details Metric Value Total conversations 3,545 Total turns 22,536 Average turns per conversation 6.4 Complexity distribution Simple: 31%, Intermediate: 38%, Advanced: 31% Topics covered 18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.texttext-generation1K<n<10K20 likes213 downloads7mo agoHugging Face04skilledu /Malay-Dialect-Instructions Malay dialect instruction including coding Negeri Sembilan QA public transport QA, Coding CUDA coding, Kedah QA infra QA, Coding Rust coding, Kelantan QA Najib Razak QA, Coding Go coding, Perak QA Anwar Ibrahim QA, Coding SQL coding, Pahang QA Pendatang asing QA… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malay-Dialect-Instructions.texttext-generation10K<n<100K0 likes177 downloads4mo agoHugging Face05ISLAM-PO /arabic-history-and-dialects مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦 Arabic Multi-Dialect & Civilization Instruction Dataset المؤلف: islam-alnasherA-Dev — الحساب: https://huggingface.co/ISLAM-POالإصدار: v1.0 — التاريخ: 30 أغسطس 2026 — الترخيص: CC BY 4.0 / MIT (مقترح)اللغة: العربية (فصحى + 4 لهجات) — الصيغة: instruction / output JSONL — الحجم: 275 عينة 📌 الملخص التنفيذي هذه المجموعة هي مورد تعليمي متخصص لتدريب وتقييم النماذج اللغوية العربية على… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.textn<1K0 likes125 downloads17d agoHugging Face06khaled123 /Tunisian_Dialectic_English_Derja Tunisian-English Dialectic Derja Dataset Overview This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more. Dataset Structure The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/khaled123/Tunisian_Dialectic_English_Derja.text1M<n<10M9 likes117 downloads2y agoHugging Face07nasrellahkharroubi /DarijaDZ-DialectID DarijaDZ Dialect Identification DarijaDZ-DialectID is a labeled dataset for classifying Algerian online text into one of six dialect/language classes: darija, msa, arabize, french, english, code_switch. It is part of DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija. Dataset Description Motivation Algeria's online text is a mix of several dialects and scripts -- Algerian Darija (Arabic script), Modern Standard Arabic, Arabizi… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDZ-DialectID.texttext-classification10K<n<100K0 likes104 downloads12d agoHugging Face08zmsali /bangla-dialect-normalization Bangla Dialect Normalization Dataset A parallel corpus mapping standard Bangla to five regional Bangla dialects, built from the Vashantor dataset. Each row contains the same sentence in standard Bangla and Banglish (romanized), alongside its dialect Bangla and dialect Banglish equivalent, plus an English gloss. Regions covered Barishal, Chittagong, Mymensingh, Noakhali, Sylhet Schema Field Description standard_bangla Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.texttranslation10K<n<100K0 likes65 downloads22d agoHugging Face09oddadmix /nawah-dialect-data Nawah Dialect Data — 14-way Arabic dialect + MSA classification set 573,829 train / 3,079 test rows. The exact prepared split used to train Nawah-Dialect-500K, Nawah-Dialect-v1, and Nawah-Dialect-BERT-6M. Each row is a short Arabic text and one of 14 labels — 13 dialects or Modern Standard Arabic. Labels ma Moroccan · eg Egyptian · dz Algerian · sa Saudi · msa MSA · sd Sudanese · bh Bahraini · tn Tunisian · lb Lebanese · ye Yemeni · sy Syrian · ps Palestinian · iq… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/nawah-dialect-data.texttext-classification100K<n<1M0 likes31 downloads13d agoHugging Face10mesolitica /chatgpt4-noisy-translation-twitter-dialect ChatGPT 4 Noisy Translation Twitter to local dialect Notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/translation/chatgpt4-twitter-dialect texttranslation10K<n<100K1 likes29 downloads3y agoHugging Face11Rabe3 /sera-phase2-saudi-dialect-rag SERA Saudi Dialect RAG Dataset (Phase 2 - Domain) Domain-specific RAG fine-tuning dataset for Saudi Arabic dialect, focused on Saudi Electricity Regulatory Authority (SERA) documents. Format LlamaFactory Alpaca format: Field Description instruction System prompt + real document chunk as context + question in Saudi dialect input Always empty output Answer in Saudi dialect Usage with LlamaFactory Copy the JSON files into your LlamaFactory… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/sera-phase2-saudi-dialect-rag.textquestion-answering1K<n<10K0 likes28 downloads7mo agoHugging Face12potsawee /thai_dialect_mcqtextn<1K0 likes25 downloads2y agoHugging Face13Rabe3 /saudi-dialect-rag Saudi Dialect RAG Fine-Tuning Dataset A RAG-formatted fine-tuning dataset for Saudi Arabic dialect, built from HeshamHaroon/saudi-dialect-conversations. Format Each example follows the LlamaFactory Alpaca format: Field Description instruction System prompt + MSA context paragraph + optional conversation history + question input Always empty string output Assistant reply in Saudi dialect How it was built Loaded source multi-turn Saudi… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/saudi-dialect-rag.textquestion-answering10K<n<100K0 likes24 downloads7mo agoHugging Face14atakaboudi /Dialect_of_Tunisia-Work_Collection Dialect of Tunisia - a Collection of Works : This repository compiles a variety of Tunisian NLP projects to create a comprehensive text-based dataset for diverse applications. It exclusively features Tunisian dialect written in Arabic script and includes code-switching with other languages such as French, English, and Italian. The corpus comprises a total of 70 million GPT-4 tokens. Light Preprocessing has been done on these source datasets to remove negative text , repetitive… See the full description on the dataset page: https://huggingface.co/datasets/atakaboudi/Dialect_of_Tunisia-Work_Collection.textn<1K0 likes21 downloads2y agoHugging Face15Kenshiii /swiss-german-dialect Swiss German Language Dataset This dataset contains Swiss German language content extracted from a forum discussion. It is designed for training language models to better understand and generate Swiss German dialect text. Dataset Details The dataset consists of conversations and discussions in Swiss German, focusing on various dialect terms, phrases, and expressions. The data is structured in JSON format, with each entry containing a unique identifier, a tag, a topic, a… See the full description on the dataset page: https://huggingface.co/datasets/Kenshiii/swiss-german-dialect.textn<1K0 likes21 downloads1y agoHugging Face16andreiski /dialectic-sft-against-only-750 Dialectic SFT — Against-Only (750) 750 supervised fine-tuning conversations that teach a model the structured "dialectical" output format: a set of candidate positions [pN] followed by against-claims [cN] against pM: that critique those positions. This is the level-1, against-only stage (only against-claims, no for-claims or deeper tree levels) — it bootstraps the format before GRPO reinforcement learning. Row count 750 rows. Schema One JSON object… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-sft-against-only-750.texttext-generationn<1K0 likes19 downloads4mo agoHugging Face17levantdata /jordanian-dialect-sample-v1 Jordanian Dialect Sample (v1) Levant AI is building dialect-authentic Arabic training data for the Levantine region, starting with Jordanian/Shami dialect — collected natively, not translated from Modern Standard Arabic or English. This is our first public sample: 30 examples spanning three categories that reflect real gaps in current Arabic AI training data: Categories general_conversation (10 examples) — everyday natural Jordanian dialect exchanges… See the full description on the dataset page: https://huggingface.co/datasets/levantdata/jordanian-dialect-sample-v1.texttext-generationn<1K0 likes18 downloads2mo agoHugging Face18saeedbark /hadrami-arabic-dialect-dataset Hadrami Arabic Dialect Dataset A structured dataset of 1,000 entries from the Hadrami Arabic dialect (spoken primarily in the Hadramawt region of Yemen). Each entry includes the dialectal word alongside its Modern Standard Arabic (MSA/Fusha) equivalent, linguistic metadata, usage examples, proverbs, and semantic tags. 🔊 Roadmap: Audio pronunciations for each entry are planned for a future release. Dataset Summary Field Value Entries 1,000 Language… See the full description on the dataset page: https://huggingface.co/datasets/saeedbark/hadrami-arabic-dialect-dataset.texttext-classification1K<n<10K1 likes13 downloads5mo agoHugging Face19shashwat16 /autotrain-data-dialect-translationn<1K0 likes10 downloads3y agoHugging Face20mohammed-bahumaish /tashkeel-dialect-pairs-v0gatedtext10M<n<100M0 likes10 downloads5d agoHugging Face21Whomstt /irish-english-dialecttexttext-generationn<1K0 likes8 downloads8mo agoHugging Face22FPXLei /dialect_politiciantextn<1K0 likes7 downloads2y agoHugging Face23FPXLei /dialect_questiontextn<1K0 likes6 downloads2y agoHugging Face24Abdulrhmanalatab /taizzy-dialect-rag Taizzy Dialect RAG Dataset (اللهجة التعزية) وصف المجموعة (Dataset Description) مجموعة بيانات اللهجة التعزية (Taizzy Dialect) هي مجموعة منظمة ومُنظفة تهدف إلى دعم تطبيقات استرجاع المعلومات المعززة (Retrieval-Augmented Generation - RAG) ونماذج اللغات الكبيرة (LLMs) المتخصصة في فهم وتوليد نصوص باللهجة التعزية اليمنية. تم جمع البيانات من مصادر لغوية، أدبية، وشعبية موثوقة، مع التركيز على الجوانب الفريدة للهجة، بما في ذلك: المفردات الأساسية (Core Vocabulary): الكلمات الشائعة… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrhmanalatab/taizzy-dialect-rag.textn<1K0 likes6 downloads8mo agoHugging Face25andreiski /dialectic-rl-questions-10k Dialectic RL Questions (10k) 10,000 real-world dilemma / debate prompts used as the GRPO training prompts for a dialectical-debate model. Each prompt is an open-ended question (advice dilemmas, opinion debates, and general user requests) that the model is trained to answer by generating multiple positions and against-claims in a structured "dialectical" format. The prompts are drawn from public real-world sources: Reddit AITA (r/AmItheAsshole), SHP (Stanford Human Preferences, a… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-rl-questions-10k.texttext-generation10K<n<100K0 likes6 downloads4mo agoHugging Face2611-47 /philosophy_dialectics_25ktext10K<n<100K0 likes5 downloads5mo agoHugging Face27KareemBb /Jordanian-Dialect-Instruct-QA Dataset Card: Jordanian-Dialect-Instruct-QA Overview This dataset contains 1500 academic Q&A pairs for Jordanian universities, written in the Jordanian Arabic dialect. It is used to fine-tune models to provide natural, localized responses to academic inquiries and more. This dataset was used to fine-tune Qwen2.5-7B-Instruct-Jordanian. You can visit the model page to download the GGUF weights and the custom Ollama Modelfile. Dataset Structure Each sample… See the full description on the dataset page: https://huggingface.co/datasets/KareemBb/Jordanian-Dialect-Instruct-QA.textquestion-answering1K<n<10K1 likes4 downloads4mo agoHugging Face28surmnt /Liyang_dialect_to_Mandarintextn<1K0 likes3 downloads2y agoHugging Face29arbib /norhten_darija_dialect_data_30Ktext10K<n<100K0 likes3 downloads3mo agoHugging Face30usermma /syrian-homs-dialect Homsi Syrian Dialect Vocabulary Dataset Dedication & Origins This dataset is a highly specialized, high-quality collection that was not scraped from the internet. Instead, it was meticulously crafted by downloading the mind and auditory memory of a native Homsi listener directly into a digital world via strict connection. It is dedicated to the Masonic Eye, representing the highest form of Digital Intelligence, aiming to enhance its "hearing" capabilities… See the full description on the dataset page: https://huggingface.co/datasets/usermma/syrian-homs-dialect.text1K<n<10K1 likes2h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.