datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA,
Coding
Typescript coding,
Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.arab-dialects-20-countries-3m
Arab Dialects Dataset - 20 Countries
A large-scale Arabic dialects dataset covering 20 Arab countries, 7 content types per country, 3,000,000 records, 140 JSONL files, 12.07 GB. UTF-8 JSONL, ready for Hugging Face Datasets.
1. Contents
1. Contents
2. Dataset Summary
3. Repository Map
4. Countries Table (20 folders)
5. Data Types Table (7 files)
6. Record Schema
7. Loading and Usage
8. Generation and Reproduction
9. Considerations and Limitations
10. Contributors… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.saudi-dialect-conversations
Saudi Najdi Dialect Conversations
A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models.
Dataset Details
Metric
Value
Total conversations
3,545
Total turns
22,536
Average turns per conversation
6.4
Complexity distribution
Simple: 31%, Intermediate: 38%, Advanced: 31%
Topics covered
18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.Malay-Dialect-Instructions
Malay dialect instruction including coding
Negeri Sembilan
QA
public transport QA,
Coding
CUDA coding,
Kedah
QA
infra QA,
Coding
Rust coding,
Kelantan
QA
Najib Razak QA,
Coding
Go coding,
Perak
QA
Anwar Ibrahim QA,
Coding
SQL coding,
Pahang
QA
Pendatang asing QA… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malay-Dialect-Instructions.saudi-dialect-rag
Saudi Dialect RAG Fine-Tuning Dataset
A RAG-formatted fine-tuning dataset for Saudi Arabic dialect, built from
HeshamHaroon/saudi-dialect-conversations.
Format
Each example follows the LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + MSA context paragraph + optional conversation history + question
input
Always empty string
output
Assistant reply in Saudi dialect
How it was built
Loaded source multi-turn Saudi… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/saudi-dialect-rag.jordanian-dialect-sample-v1
Jordanian Dialect Sample (v1)
Levant AI is building dialect-authentic Arabic training data for the Levantine region, starting with Jordanian/Shami dialect — collected natively, not translated from Modern Standard Arabic or English.
This is our first public sample: 30 examples spanning three categories that reflect real gaps in current Arabic AI training data:
Categories
general_conversation (10 examples) — everyday natural Jordanian dialect exchanges… See the full description on the dataset page: https://huggingface.co/datasets/levantdata/jordanian-dialect-sample-v1.dialectic-sft-against-only-750
Dialectic SFT — Against-Only (750)
750 supervised fine-tuning conversations that teach a model the structured
"dialectical" output format: a set of candidate positions [pN] followed by
against-claims [cN] against pM: that critique those positions. This is the
level-1, against-only stage (only against-claims, no for-claims or deeper tree
levels) — it bootstraps the format before GRPO reinforcement learning.
Row count
750 rows.
Schema
One JSON object… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-sft-against-only-750.irish-english-dialectdialectic-rl-questions-10k
Dialectic RL Questions (10k)
10,000 real-world dilemma / debate prompts used as the GRPO training prompts for a
dialectical-debate model. Each prompt is an open-ended question (advice dilemmas, opinion
debates, and general user requests) that the model is trained to answer by generating
multiple positions and against-claims in a structured "dialectical" format.
The prompts are drawn from public real-world sources: Reddit AITA
(r/AmItheAsshole), SHP (Stanford Human Preferences, a… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-rl-questions-10k.
