datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SADA22
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of transcribed Arabic audio recordings, primarily featuring various Saudi dialects, and was curated in a collaboration between the National Center for Artificial… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/SADA22.sada-validation-preprocessed
Details
This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same tokenizer. Please double check this as well from the original model repo.
In addtition, the following filters were applied to this data:
All audios are less than 30 seconds and greater than 0… See the full description on the dataset page: https://huggingface.co/datasets/mosama/sada-validation-preprocessed.arabic-speech-SADA22-MSA
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
⚠️ Caution
This is only the portion of the SADA dataset where the speaker dialect is Modern Standard Arabic (MSA). To access full dataset, you should check this link.
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-MSA.SADA-Najdi-from-kaggle
SADA 2022 — Najdi Dialect (Preprocessed for TTS)
Najdi dialect subset of SADA 2022,
preprocessed and ready for XTTS-v2 fine-tuning.
Preprocessing Pipeline
SADA full episodes → Filter Najdi → Slice by SegmentStart/End
→ Resample 22050Hz → Trim silence → Peak normalize
→ Duration filter (2.0–11.0s)
→ SNR filter (≥12.0dB)
→ Environment filter (Clean only)
→ Arabic text normalization
Step
Details
Dialect
SpeakerDialect == "Najdi"
Environment… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed007/SADA-Najdi-from-kaggle.arabic-speech-SADA22-Khaliji
Dataset Card for SADA (Saudi Audio Dataset for Arabic)
⚠️ Caution
This is only the portion of the SADA dataset where the speaker dialect is Khaliji. To access full dataset, you should check this link.
Dataset Summary
The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-Khaliji.sada-diarization-preview
SADA 2022 Arabic Diarization
Training-ready speaker-attributed ASR windows derived from
SADA 2022. The source
recordings are mirrored at
khaledalganem/sada2022.
Splits
train: 36,004 windows, 202.064 hours, 4,062 recordings
validation: 853 windows, 4.774 hours, 88 recordings
test: 901 windows, 5.006 hours, 111 recordings
Total: 37,758 windows and
211.844 hours.
The official SADA train, validation, and test partitions are preserved.
Windows are 8–28 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Mohaddz/sada-diarization-preview.Clean_One_Speaker_SADA
Clean_One_Speaker_SADA
A cleaned, single-speaker subset of the SADA (Saudi Audio Dataset for Arabic) corpus,
derived from MahmoudIbrahim/100Hours-SADA22.
Processing
Follows the cleaning procedure from the Kaggle notebook
Segmented Audio Data for Arabic Dialects (SADA):
Removed rows with Unknown speaker age or gender.
Removed rows whose dialect is More than 1 speaker, Unknown, or Notapplicable
(every remaining segment has exactly one identified speaker).… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_SADA.Clean_One_Speaker_SADA_Split
Clean_One_Speaker_SADA_Split
Train/validation/test version of
Sebssihakim/Clean_One_Speaker_SADA,
a cleaned single-speaker subset of the SADA corpus (SDAIA / Saudi Broadcasting Authority).
Splits
80/10/10, stratified by speaker_dialect, seed 42:
train ~30.8k, validation ~3.9k, test ~3.9k rows.
The Maghrebi dialect (7 rows) was removed — too few samples to stratify.
Splits are segment-level: the same source show/speaker may appear in more than one split.
Rare… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_SADA_Split.
