datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
saudi-dialect-conversations
Saudi Najdi Dialect Conversations
A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models.
Dataset Details
Metric
Value
Total conversations
3,545
Total turns
22,536
Average turns per conversation
6.4
Complexity distribution
Simple: 31%, Intermediate: 38%, Advanced: 31%
Topics covered
18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.saudi_dialect_asrv1.0saudi-dialect-speech-female
🌍 Saudi Dialectal Arabic Audio Dataset
This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning.
🗂️ Dataset Columns
Column
Description
audio
The audio chunk (22,050 Hz, mono WAV)
duration
Chunk duration in seconds
base_transcription
Transcript from the base Arabic ASR model
dialectal_transcription
Transcript from the Saudi-dialectal… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/saudi-dialect-speech-female.saudi-dialect-speech-malesaudi-dialect-test-samples
Saudi Dialect Test Samples
Dataset Description
This dataset contains 1280 Saudi dialect utterances across 44 categories, used for testing and evaluating the Omartificial-Intelligence-Space/SA-BERT-V1 model. The sentences represent a wide range of topics, from daily conversations to specialized domains.
Dataset Structure
Data Fields
category: The topic category of the utterance (one of 44 categories)
text: The Saudi dialect text… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/saudi-dialect-test-samples.SaudiDialect-Triplet-21
📂 SaudiDialect-Triplet-21 : Saudi Triplet Dataset (SABER Training Data)
🧩 Dataset Summary
The Saudi Triplet Dataset is a high-quality corpus of 2,964 sentence triplets (Anchor, Positive, Negative) specifically curated to capture the nuances of Saudi Arabic dialects (Najdi, Hijazi, Gulf, etc.).
This dataset was created to fine-tune semantic embedding models such as SABER for tasks like Semantic Search, Retrieval-Augmented Generation (RAG), and Clustering.
It covers 21… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/SaudiDialect-Triplet-21.arabic-eou-saudi-dialect
Arabic End-of-Utterance Detection Dataset (Saudi Dialect)
Dataset Description
This dataset is designed for training and evaluating End-of-Utterance (EOU) detection models for Arabic conversations, with emphasis on Saudi dialect patterns.
Dataset Summary
The dataset contains Arabic conversational samples labeled for binary classification:
Positive (1): End of utterance - speaker has finished their turn
Negative (0): Not end of utterance - speaker will continue… See the full description on the dataset page: https://huggingface.co/datasets/mahmoudsaalama/arabic-eou-saudi-dialect.sera-phase2-saudi-dialect-rag
SERA Saudi Dialect RAG Dataset (Phase 2 - Domain)
Domain-specific RAG fine-tuning dataset for Saudi Arabic dialect,
focused on Saudi Electricity Regulatory Authority (SERA) documents.
Format
LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + real document chunk as context + question in Saudi dialect
input
Always empty
output
Answer in Saudi dialect
Usage with LlamaFactory
Copy the JSON files into your LlamaFactory… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/sera-phase2-saudi-dialect-rag.saudi-dialect-rag
Saudi Dialect RAG Fine-Tuning Dataset
A RAG-formatted fine-tuning dataset for Saudi Arabic dialect, built from
HeshamHaroon/saudi-dialect-conversations.
Format
Each example follows the LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + MSA context paragraph + optional conversation history + question
input
Always empty string
output
Assistant reply in Saudi dialect
How it was built
Loaded source multi-turn Saudi… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/saudi-dialect-rag.sada-eou-saudi-dialect
