datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA,
Coding
Typescript coding,
Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.arab-dialects-20-countries-3m
Arab Dialects Dataset - 20 Countries
A large-scale Arabic dialects dataset covering 20 Arab countries, 7 content types per country, 3,000,000 records, 140 JSONL files, 12.07 GB. UTF-8 JSONL, ready for Hugging Face Datasets.
1. Contents
1. Contents
2. Dataset Summary
3. Repository Map
4. Countries Table (20 folders)
5. Data Types Table (7 files)
6. Record Schema
7. Loading and Usage
8. Generation and Reproduction
9. Considerations and Limitations
10. Contributors… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.dialect-preferences
DiaLLM — Pooled Preference Dataset (Implicit Thread)
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in
English Dialect Adaptation (EMNLP 2026 Main).
45,690 preference pairs, pooling all three variety-specific sets
(Australian,
Northern British,
Indian) without
variety targeting. Used for implicit-thread DPO training, where the three
varieties are pooled rather than targeted individually, preserving the
variety-agnostic objective of that thread.… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/dialect-preferences.saudi-dialect-conversations
Saudi Najdi Dialect Conversations
A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models.
Dataset Details
Metric
Value
Total conversations
3,545
Total turns
22,536
Average turns per conversation
6.4
Complexity distribution
Simple: 31%, Intermediate: 38%, Advanced: 31%
Topics covered
18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.shironaam
Dataset Card for Shironaam Corpus
Dataset Summary
Automatic headline generation systems have the potential to assist editors in finding interesting headlines to attract visitors or readers.
However, the performance of headline generation systems remains challenging due to the unavailability of sufficient parallel data for
low-resource languages like Bengali. We provide Shironaam, a large-scale news headline generation dataset of a low-resource language
i.e., Bengali… See the full description on the dataset page: https://huggingface.co/datasets/dialect-ai/shironaam.Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malay-Dialect-Instructions.arabic-dialect-corpus
Arabic Dialect Corpus
A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata.
Dataset Statistics
Metric
Value
Total Records
127,180
Total Tokens
5,802,324
Average Tokens per Record
45.62
Dialect Categories
5
Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.Dialect2SQL
Dialect2SQL
Dataset Description
Dialect2SQL is a novel dataset designed for the Text-to-SQL task in Arabic dialects, with a particular focus on Moroccan Darija.It provides natural language questions written in Darija, paired with corresponding SQL queries and database schemas.The dataset enables research on low-resource natural language interfaces to databases (NLIDB) in non-standard Arabic varieties.
Dataset Summary
Dialect2SQL aims to bridge the gap… See the full description on the dataset page: https://huggingface.co/datasets/salmane11/Dialect2SQL.arabic-dialect-corpus
🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi)
Dataset Description
This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA).
Languages
Primary Dialects:
Egyptian Arabic (EG) - Cairene and regional Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/arabic-dialect-corpus.swiss-dialects
Dataset Card for ArchiMod Corpus
Dataset Summary
The ArchiMob corpus represents German linguistic varieties spoken within the territory of Switzerland. This corpus is the first electronic resource containing long samples of transcribed text in Swiss German, intended for studying the spatial distribution of morphosyntactic features and for natural language processing.
Languages
Swiss-German
Dataset Structure
Data Instances
{ 'sentence':… See the full description on the dataset page: https://huggingface.co/datasets/statworx/swiss-dialects.dialectical-reasoningA specialised dialectical reasoning dataset.
contain { Thesis:, Antithesis:, Synthesis: }.
Domain are math, science, creative writing
arabic-dialect-dpo
Arabic Dialect DPO Dataset - Egyptian & Saudi
The first large-scale Arabic dialect preference dataset for DPO/ORPO/GRPO alignment training. Contains 22,538 preference triples across two major Arabic dialects: Egyptian (Masry) and Saudi (Najdi).
Dataset Summary
Config
Dialect
Rows
Language Code
egyptian
Egyptian Arabic (مصري)
11,038
ar-EG
saudi
Saudi Arabic (سعودي نجدي)
11,500
ar-SA
Total
22,538
Usage
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-dialect-dpo.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.arabic-dialect-corpus
🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi)
Dataset Description
This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA).
Languages
Primary Dialects:
Egyptian Arabic (EG) - Cairene and regional… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/arabic-dialect-corpus.dialect-at-tirol
Dataset of Tyrolean Dialect (Austria)
This dataset contains 200+ words used in Tirol (Austria), together with their German translation and (optional) meaning.
TheArabicPile_Dialects
The Arabic Pile
Introduction:
The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Dialects.saudi-dialect-rag
Saudi Dialect RAG Fine-Tuning Dataset
A RAG-formatted fine-tuning dataset for Saudi Arabic dialect, built from
HeshamHaroon/saudi-dialect-conversations.
Format
Each example follows the LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + MSA context paragraph + optional conversation history + question
input
Always empty string
output
Assistant reply in Saudi dialect
How it was built
Loaded source multi-turn Saudi… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/saudi-dialect-rag.jordanian-dialect-sample-v1
Jordanian Dialect Sample (v1)
Levant AI is building dialect-authentic Arabic training data for the Levantine region, starting with Jordanian/Shami dialect — collected natively, not translated from Modern Standard Arabic or English.
This is our first public sample: 30 examples spanning three categories that reflect real gaps in current Arabic AI training data:
Categories
general_conversation (10 examples) — everyday natural Jordanian dialect exchanges… See the full description on the dataset page: https://huggingface.co/datasets/levantdata/jordanian-dialect-sample-v1.dialectic-preferences-bias-aae-sae-parallel
Dialectic Preferences Bias Dataset
Dataset Description
Overview
This dataset is part of a research study examining dialectic preference bias in Large Language Models (LLMs). It contains paired sentences in African American English (AAE) and Standard American English (SAE), used to analyze potential biases in language models' treatment of different dialects.
The dataset contains two columns:
african_american_english: Text samples in African American English… See the full description on the dataset page: https://huggingface.co/datasets/furquan/dialectic-preferences-bias-aae-sae-parallel.dialectic-reasoning-traces
Dialectic Reasoning Traces
255 scored dialectic reasoning traces for training models on integrative resolution under conflicting frames. Instead of list-format pros/cons or generic hedging, these traces teach models to identify real tension, make conditional commitments, and reach specific resolutions.
Version Note
This dataset contains only v1 traces — the clean, non-fabricating training data. An earlier version on this repo included augmented data from later pipeline… See the full description on the dataset page: https://huggingface.co/datasets/hikewa/dialectic-reasoning-traces.dialectic-sft-against-only-750
Dialectic SFT — Against-Only (750)
750 supervised fine-tuning conversations that teach a model the structured
"dialectical" output format: a set of candidate positions [pN] followed by
against-claims [cN] against pM: that critique those positions. This is the
level-1, against-only stage (only against-claims, no for-claims or deeper tree
levels) — it bootstraps the format before GRPO reinforcement learning.
Row count
750 rows.
Schema
One JSON object… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-sft-against-only-750.Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT
Turkish Dialectical Reasoning Dataset (Sokrates-ToT)
The Turkish Dialectical Reasoning Dataset (Sokrates-ToT) is a collection structured in a Tree-of-Thought (ToT) format, based on a multi-persona and dialectical reasoning framework.Inspired by Socrates' method of dialogue, it facilitates deep analysis of complex and multidimensional issues by having AI personas with different expertise interact and ultimately reach a final synthesis.
Purpose of the Dataset
This… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT.irish-english-dialectUnderstanding_Dialect_Text
🇰🇿 Kazakh Dialect Analysis and Standardization Dataset
Dataset Summary
Kazakh Dialect Analysis and Standardization Dataset is a Kazakh-language linguistics dataset designed for dialect identification, dialectal feature analysis, and normalization into standard Kazakh.
The dataset contains instruction-style prompts asking the model to identify dialectal or regional language features in a given Kazakh text and explain how they can be converted into standard… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Understanding_Dialect_Text.dialectic-rl-questions-10k
Dialectic RL Questions (10k)
10,000 real-world dilemma / debate prompts used as the GRPO training prompts for a
dialectical-debate model. Each prompt is an open-ended question (advice dilemmas, opinion
debates, and general user requests) that the model is trained to answer by generating
multiple positions and against-claims in a structured "dialectical" format.
The prompts are drawn from public real-world sources: Reddit AITA
(r/AmItheAsshole), SHP (Stanford Human Preferences, a… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-rl-questions-10k.
