oddadmix/arabic-audio-collection-sudanese-ahmed-gobara
Ahmed Gobara Sudanese Arabic Speech Dataset Dataset Summary The Ahmed Gobara Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 19 hours of speech recordings and corresponding transcripts. While compact, the dataset offers a clean, consistent single-speaker resource in Sudanese Arabic — an Arabic variety with very few open speech resources — making it especially valuable for voice cloning, speaker adaptation… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-ahmed-gobara.
Ahmed Gobara Sudanese Arabic Speech Dataset
Dataset Summary
The Ahmed Gobara Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 19 hours of speech recordings and corresponding transcripts.
While compact, the dataset offers a clean, consistent single-speaker resource in Sudanese Arabic — an Arabic variety with very few open speech resources — making it especially valuable for voice cloning, speaker adaptation, and dialect-focused fine-tuning experiments.
Alongside the spoken Arabic text, the transcripts capture paralinguistic vocalizations, breathing, and emotional cues, enabling the development of expressive, natural-sounding AI models that go beyond flat dictation.
The dataset was created to support Arabic speech technology research and development, including:
- Automatic Speech Recognition (ASR)
- Text-to-Speech (TTS) and Expressive TTS
- Speech Foundation Models
- Audio-Text Alignment
- Speaker Adaptation
- Paralinguistic and Emotion Recognition
- Sudanese Dialect and Arabic Language Technology Research
Transcriptions were generated and curated by the dataset creator using an AI-assisted transcription pipeline and additional quality-control procedures to ensure the accurate logging of non-verbal tokens.
Non-Verbal Vocalization Tags
The transcripts include a comprehensive set of non-verbal tokens to capture the true nuance of human speech and breathing. The supported tags are:
<laugh> <cry> <weep> <sob> <scream> <shout> <whisper> <sigh> <gasp> <groan> <moan> <pause> <hes> <stutter> <breath> <sniff> <cough> <throat_clear>
Research & Application Use Cases
With ~19 hours of single-speaker Sudanese Arabic speech and granular non-verbal tagging, the dataset is particularly suitable for:
- Sudanese Arabic TTS and voice cloning: A consistent single-speaker voice for dialect-aware, expressive synthesis.
- Speaker adaptation: Fine-tuning existing Arabic ASR/TTS models to a Sudanese speaker with limited data.
- ASR fine-tuning and evaluation for Sudanese Arabic.
- Foundation model fine-tuning.
- Research on Sudanese Arabic speech, narration styles, and paralinguistics.
Dataset Statistics
يحتفظ المنشئون الأصليون والقنوات المالكة بكافة الحقوق، ولا يتم ادعاء أي ملكية للملفات الصوتية الأصلية أو المحتوى الصوتي الأساسي.
