oddadmix/arabic-audio-collection-mostafa-mahmoud
Mostafa Mahmoud Arabic Speech Dataset Dataset Summary The Mostafa Mahmoud Arabic Speech Dataset is a large-scale Arabic speech corpus containing approximately 187 hours of speech recordings and corresponding transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud. The dataset was created to support Arabic speech technology research and development, including: Automatic Speech Recognition (ASR)… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mostafa-mahmoud.
Mostafa Mahmoud Arabic Speech Dataset
Dataset Summary
The Mostafa Mahmoud Arabic Speech Dataset is a large-scale Arabic speech corpus containing approximately 187 hours of speech recordings and corresponding transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud.
The dataset was created to support Arabic speech technology research and development, including:
- Automatic Speech Recognition (ASR)
- Text-to-Speech (TTS)
- Speech Foundation Models
- Audio-Text Alignment
- Speaker Adaptation
- Arabic Language Technology Research
Transcriptions were generated and curated by the dataset creator using an AI-assisted transcription pipeline and additional quality-control procedures.
With 187 hours of speech from a single speaker, the dataset is particularly suitable for:
- High-quality Arabic TTS voice cloning
- Speaker adaptation and speaker representation learning
- Long-form ASR training
- Foundation model pretraining and fine-tuning
- Research on Arabic speech and narration styles
Dataset Statistics
يحتفظ المنشئون الأصليون والقنوات المالكة بكافة الحقوق، ولا يتم ادعاء أي ملكية للملفات الصوتية الأصلية أو المحتوى الصوتي الأساسي.
