datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
saudi-dialect-conversations
Saudi Najdi Dialect Conversations
A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models.
Dataset Details
Metric
Value
Total conversations
3,545
Total turns
22,536
Average turns per conversation
6.4
Complexity distribution
Simple: 31%, Intermediate: 38%, Advanced: 31%
Topics covered
18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.sera-phase2-saudi-dialect-rag
SERA Saudi Dialect RAG Dataset (Phase 2 - Domain)
Domain-specific RAG fine-tuning dataset for Saudi Arabic dialect,
focused on Saudi Electricity Regulatory Authority (SERA) documents.
Format
LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + real document chunk as context + question in Saudi dialect
input
Always empty
output
Answer in Saudi dialect
Usage with LlamaFactory
Copy the JSON files into your LlamaFactory… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/sera-phase2-saudi-dialect-rag.saudi-dialect-rag
Saudi Dialect RAG Fine-Tuning Dataset
A RAG-formatted fine-tuning dataset for Saudi Arabic dialect, built from
HeshamHaroon/saudi-dialect-conversations.
Format
Each example follows the LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + MSA context paragraph + optional conversation history + question
input
Always empty string
output
Assistant reply in Saudi dialect
How it was built
Loaded source multi-turn Saudi… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/saudi-dialect-rag.
