datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sada22-najdinazrah-najdi-voice-datasetFineWeb2-Najdi-Arabic
FineWeb2 Najdi Arabic
🇸🇦 This is the Nadj Arabic Portion of The FineWeb2 Dataset.
🇸🇦 This dataset contains a comprehensive collection of text in Najdi Arabic, a regional dialect within the Afro-Asiatic language family. With over 562 million words and 1.6 million documents, it provides a valuable resource for developing NLP tools and applications specific to Najdi Arabic.
Purpose of This Repository
This repository provides easy access to the Arabic portion - Najdi of… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-Najdi-Arabic.youtube_najdiSADA-Najdi-from-kaggle
SADA 2022 — Najdi Dialect (Preprocessed for TTS)
Najdi dialect subset of SADA 2022,
preprocessed and ready for XTTS-v2 fine-tuning.
Preprocessing Pipeline
SADA full episodes → Filter Najdi → Slice by SegmentStart/End
→ Resample 22050Hz → Trim silence → Peak normalize
→ Duration filter (2.0–11.0s)
→ SNR filter (≥12.0dB)
→ Environment filter (Clean only)
→ Arabic text normalization
Step
Details
Dialect
SpeakerDialect == "Najdi"
Environment… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed007/SADA-Najdi-from-kaggle.najdi_mt
saleh1312/najdi_mt
Spoken Najdi Hadari conversations in OpenAI multi-turn chat format:
{"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Mix
source
label
count
percent
convs
محادثات عادية
66
25.7%
word_meanings
معاني الكلمات
20
7.8%
word_usages
استخدام طبيعي
129
50.2%
word_multiturn
محادثات متعددة الأدوار
42
16.3%
total
257
100%
sada2022-Najdi-CleanFineTranslations_Najdinajdi-female-tts-small
Dataset Card for SADA - Saudi Audio Dataset for Arabic
Dataset Details
Dataset Description
The SADA (Saudi Audio Dataset for Arabic) is a comprehensive dataset consisting of audio recordings from over 57 TV shows aired by the Saudi Broadcasting Authority (SBA). The dataset contains approximately 667 hours of audio data with transcripts, the majority of which are in various Saudi dialects (Najdi, Hijazi, Khaliji, etc.).
Curated by: The National Center for… See the full description on the dataset page: https://huggingface.co/datasets/muaadh019/najdi-female-tts-small.sada-najdi-validation-ttsnajdi-female-tts-smalll
