datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sada22-najdinazrah-najdi-voice-datasetyoutube_najdiSADA-Najdi-from-kaggle
SADA 2022 — Najdi Dialect (Preprocessed for TTS)
Najdi dialect subset of SADA 2022,
preprocessed and ready for XTTS-v2 fine-tuning.
Preprocessing Pipeline
SADA full episodes → Filter Najdi → Slice by SegmentStart/End
→ Resample 22050Hz → Trim silence → Peak normalize
→ Duration filter (2.0–11.0s)
→ SNR filter (≥12.0dB)
→ Environment filter (Clean only)
→ Arabic text normalization
Step
Details
Dialect
SpeakerDialect == "Najdi"
Environment… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed007/SADA-Najdi-from-kaggle.sada2022-Najdi-Cleannajdi-female-tts-small
Dataset Card for SADA - Saudi Audio Dataset for Arabic
Dataset Details
Dataset Description
The SADA (Saudi Audio Dataset for Arabic) is a comprehensive dataset consisting of audio recordings from over 57 TV shows aired by the Saudi Broadcasting Authority (SBA). The dataset contains approximately 667 hours of audio data with transcripts, the majority of which are in various Saudi dialects (Najdi, Hijazi, Khaliji, etc.).
Curated by: The National Center for… See the full description on the dataset page: https://huggingface.co/datasets/muaadh019/najdi-female-tts-small.sada-najdi-validation-tts
