datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.arabic-tashkeel-speech
Nahw Arabic Tashkeel Speech Dataset
An open-source collection of 1,093 fully diacritized Arabic speech recordings, crowd-sourced from native speakers via Nahw.ai.
Dataset summary
Stat
Value
Total recordings
1,093
Speakers
10
Language
Arabic (ar)
Sampling rate
16 kHz
License
CC-BY-4.0
Features
audio: The speech recording, resampled to 16 kHz.
transcription: The fully diacritized Arabic sentence that was read aloud.
sentence: The same… See the full description on the dataset page: https://huggingface.co/datasets/NahwAI/arabic-tashkeel-speech.tashkeelal_sallom_UAE_transcription_by_elevenlab_tashkeeltashkeela-audio
