Saudi Dialect
saudi-dialect-conversations
Saudi Najdi Dialect Conversations
A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models.
Dataset Details
Metric
Value
Total conversations
3,545
Total turns
22,536
Average turns per conversation
6.4
Complexity distribution
Simple: 31%, Intermediate: 38%, Advanced: 31%
Topics covered
18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.saudi_dialect_asrv1.0saudi-dialect-speech-female
🌍 Saudi Dialectal Arabic Audio Dataset
This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning.
🗂️ Dataset Columns
Column
Description
audio
The audio chunk (22,050 Hz, mono WAV)
duration
Chunk duration in seconds
base_transcription
Transcript from the base Arabic ASR model
dialectal_transcription
Transcript from the Saudi-dialectal… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/saudi-dialect-speech-female.saudi-dialect-speech-malesaudi-dialect-test-samples
Saudi Dialect Test Samples
Dataset Description
This dataset contains 1280 Saudi dialect utterances across 44 categories, used for testing and evaluating the Omartificial-Intelligence-Space/SA-BERT-V1 model. The sentences represent a wide range of topics, from daily conversations to specialized domains.
Dataset Structure
Data Fields
category: The topic category of the utterance (one of 44 categories)
text: The Saudi dialect text… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/saudi-dialect-test-samples.SaudiDialect-Triplet-21
📂 SaudiDialect-Triplet-21 : Saudi Triplet Dataset (SABER Training Data)
🧩 Dataset Summary
The Saudi Triplet Dataset is a high-quality corpus of 2,964 sentence triplets (Anchor, Positive, Negative) specifically curated to capture the nuances of Saudi Arabic dialects (Najdi, Hijazi, Gulf, etc.).
This dataset was created to fine-tune semantic embedding models such as SABER for tasks like Semantic Search, Retrieval-Augmented Generation (RAG), and Clustering.
It covers 21… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/SaudiDialect-Triplet-21.
