datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA,
Coding
Typescript coding,
Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.arab-dialects-20-countries-3m
Arab Dialects Dataset - 20 Countries
A large-scale Arabic dialects dataset covering 20 Arab countries, 7 content types per country, 3,000,000 records, 140 JSONL files, 12.07 GB. UTF-8 JSONL, ready for Hugging Face Datasets.
1. Contents
1. Contents
2. Dataset Summary
3. Repository Map
4. Countries Table (20 folders)
5. Data Types Table (7 files)
6. Record Schema
7. Loading and Usage
8. Generation and Reproduction
9. Considerations and Limitations
10. Contributors… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.saudi-dialect-conversations
Saudi Najdi Dialect Conversations
A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models.
Dataset Details
Metric
Value
Total conversations
3,545
Total turns
22,536
Average turns per conversation
6.4
Complexity distribution
Simple: 31%, Intermediate: 38%, Advanced: 31%
Topics covered
18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malay-Dialect-Instructions.arabic-history-and-dialects
مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦
Arabic Multi-Dialect & Civilization Instruction Dataset
المؤلف: islam-alnasherA-Dev — الحساب: https://huggingface.co/ISLAM-POالإصدار: v1.0 — التاريخ: 30 أغسطس 2026 — الترخيص: CC BY 4.0 / MIT (مقترح)اللغة: العربية (فصحى + 4 لهجات) — الصيغة: instruction / output JSONL — الحجم: 275 عينة
📌 الملخص التنفيذي
هذه المجموعة هي مورد تعليمي متخصص لتدريب وتقييم النماذج اللغوية العربية على… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.Tunisian_Dialectic_English_Derja
Tunisian-English Dialectic Derja Dataset
Overview
This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more.
Dataset Structure
The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/khaled123/Tunisian_Dialectic_English_Derja.DarijaDZ-DialectID
DarijaDZ Dialect Identification
DarijaDZ-DialectID is a labeled dataset for classifying Algerian
online text into one of six dialect/language classes: darija, msa,
arabize, french, english, code_switch. It is part of
DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija.
Dataset Description
Motivation
Algeria's online text is a mix of several dialects and scripts --
Algerian Darija (Arabic script), Modern Standard Arabic, Arabizi… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDZ-DialectID.bangla-dialect-normalization
Bangla Dialect Normalization Dataset
A parallel corpus mapping standard Bangla to five regional Bangla dialects,
built from the Vashantor dataset. Each row contains the same sentence in
standard Bangla and Banglish (romanized), alongside its dialect Bangla and
dialect Banglish equivalent, plus an English gloss.
Regions covered
Barishal, Chittagong, Mymensingh, Noakhali, Sylhet
Schema
Field
Description
standard_bangla
Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.nawah-dialect-data
Nawah Dialect Data — 14-way Arabic dialect + MSA classification set
573,829 train / 3,079 test rows. The exact prepared split used to train
Nawah-Dialect-500K,
Nawah-Dialect-v1, and
Nawah-Dialect-BERT-6M. Each row is a
short Arabic text and one of 14 labels — 13 dialects or Modern Standard Arabic.
Labels
ma Moroccan · eg Egyptian · dz Algerian · sa Saudi · msa MSA · sd Sudanese ·
bh Bahraini · tn Tunisian · lb Lebanese · ye Yemeni · sy Syrian · ps Palestinian ·
iq… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/nawah-dialect-data.chatgpt4-noisy-translation-twitter-dialect
ChatGPT 4 Noisy Translation Twitter to local dialect
Notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/translation/chatgpt4-twitter-dialect
sera-phase2-saudi-dialect-rag
SERA Saudi Dialect RAG Dataset (Phase 2 - Domain)
Domain-specific RAG fine-tuning dataset for Saudi Arabic dialect,
focused on Saudi Electricity Regulatory Authority (SERA) documents.
Format
LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + real document chunk as context + question in Saudi dialect
input
Always empty
output
Answer in Saudi dialect
Usage with LlamaFactory
Copy the JSON files into your LlamaFactory… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/sera-phase2-saudi-dialect-rag.thai_dialect_mcqsaudi-dialect-rag
Saudi Dialect RAG Fine-Tuning Dataset
A RAG-formatted fine-tuning dataset for Saudi Arabic dialect, built from
HeshamHaroon/saudi-dialect-conversations.
Format
Each example follows the LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + MSA context paragraph + optional conversation history + question
input
Always empty string
output
Assistant reply in Saudi dialect
How it was built
Loaded source multi-turn Saudi… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/saudi-dialect-rag.Dialect_of_Tunisia-Work_Collection
Dialect of Tunisia - a Collection of Works :
This repository compiles a variety of Tunisian NLP projects to create a comprehensive text-based dataset for diverse applications. It exclusively features Tunisian dialect written in Arabic script and includes code-switching with other languages such as French, English, and Italian. The corpus comprises a total of 70 million GPT-4 tokens. Light Preprocessing has been done on these source datasets to remove negative text , repetitive… See the full description on the dataset page: https://huggingface.co/datasets/atakaboudi/Dialect_of_Tunisia-Work_Collection.swiss-german-dialect
Swiss German Language Dataset
This dataset contains Swiss German language content extracted from a forum discussion. It is designed for training language models to better understand and generate Swiss German dialect text.
Dataset Details
The dataset consists of conversations and discussions in Swiss German, focusing on various dialect terms, phrases, and expressions. The data is structured in JSON format, with each entry containing a unique identifier, a tag, a topic, a… See the full description on the dataset page: https://huggingface.co/datasets/Kenshiii/swiss-german-dialect.dialectic-sft-against-only-750
Dialectic SFT — Against-Only (750)
750 supervised fine-tuning conversations that teach a model the structured
"dialectical" output format: a set of candidate positions [pN] followed by
against-claims [cN] against pM: that critique those positions. This is the
level-1, against-only stage (only against-claims, no for-claims or deeper tree
levels) — it bootstraps the format before GRPO reinforcement learning.
Row count
750 rows.
Schema
One JSON object… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-sft-against-only-750.jordanian-dialect-sample-v1
Jordanian Dialect Sample (v1)
Levant AI is building dialect-authentic Arabic training data for the Levantine region, starting with Jordanian/Shami dialect — collected natively, not translated from Modern Standard Arabic or English.
This is our first public sample: 30 examples spanning three categories that reflect real gaps in current Arabic AI training data:
Categories
general_conversation (10 examples) — everyday natural Jordanian dialect exchanges… See the full description on the dataset page: https://huggingface.co/datasets/levantdata/jordanian-dialect-sample-v1.hadrami-arabic-dialect-dataset
Hadrami Arabic Dialect Dataset
A structured dataset of 1,000 entries from the Hadrami Arabic dialect (spoken primarily in the Hadramawt region of Yemen). Each entry includes the dialectal word alongside its Modern Standard Arabic (MSA/Fusha) equivalent, linguistic metadata, usage examples, proverbs, and semantic tags.
🔊 Roadmap: Audio pronunciations for each entry are planned for a future release.
Dataset Summary
Field
Value
Entries
1,000
Language… See the full description on the dataset page: https://huggingface.co/datasets/saeedbark/hadrami-arabic-dialect-dataset.autotrain-data-dialect-translationtashkeel-dialect-pairs-v0irish-english-dialectdialect_politiciandialect_questiontaizzy-dialect-rag
Taizzy Dialect RAG Dataset (اللهجة التعزية)
وصف المجموعة (Dataset Description)
مجموعة بيانات اللهجة التعزية (Taizzy Dialect) هي مجموعة منظمة ومُنظفة تهدف إلى دعم تطبيقات استرجاع المعلومات المعززة (Retrieval-Augmented Generation - RAG) ونماذج اللغات الكبيرة (LLMs) المتخصصة في فهم وتوليد نصوص باللهجة التعزية اليمنية.
تم جمع البيانات من مصادر لغوية، أدبية، وشعبية موثوقة، مع التركيز على الجوانب الفريدة للهجة، بما في ذلك:
المفردات الأساسية (Core Vocabulary): الكلمات الشائعة… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrhmanalatab/taizzy-dialect-rag.dialectic-rl-questions-10k
Dialectic RL Questions (10k)
10,000 real-world dilemma / debate prompts used as the GRPO training prompts for a
dialectical-debate model. Each prompt is an open-ended question (advice dilemmas, opinion
debates, and general user requests) that the model is trained to answer by generating
multiple positions and against-claims in a structured "dialectical" format.
The prompts are drawn from public real-world sources: Reddit AITA
(r/AmItheAsshole), SHP (Stanford Human Preferences, a… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-rl-questions-10k.philosophy_dialectics_25kJordanian-Dialect-Instruct-QA
Dataset Card: Jordanian-Dialect-Instruct-QA
Overview
This dataset contains 1500 academic Q&A pairs for Jordanian universities, written in the Jordanian Arabic dialect. It is used to fine-tune models to provide natural, localized responses to academic inquiries and more.
This dataset was used to fine-tune Qwen2.5-7B-Instruct-Jordanian. You can visit the model page to download the GGUF weights and the custom Ollama Modelfile.
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/KareemBb/Jordanian-Dialect-Instruct-QA.Liyang_dialect_to_Mandarinnorhten_darija_dialect_data_30Ksyrian-homs-dialect
Homsi Syrian Dialect Vocabulary Dataset
Dedication & Origins
This dataset is a highly specialized, high-quality collection that was not scraped from the internet. Instead, it was meticulously crafted by downloading the mind and auditory memory of a native Homsi listener directly into a digital world via strict connection.
It is dedicated to the Masonic Eye, representing the highest form of Digital Intelligence, aiming to enhance its "hearing" capabilities… See the full description on the dataset page: https://huggingface.co/datasets/usermma/syrian-homs-dialect.
