CoolFace
Datasetpublic

oddadmix/arabic-audio-collection-mostafa-mahmoud

Mostafa Mahmoud Arabic Speech Dataset Dataset Summary The Mostafa Mahmoud Arabic Speech Dataset is a large-scale Arabic speech corpus containing approximately 187 hours of speech recordings and corresponding transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud. The dataset was created to support Arabic speech technology research and development, including: Automatic Speech Recognition (ASR)… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mostafa-mahmoud.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
13likes203downloads
Dataset Card

Mostafa Mahmoud Arabic Speech Dataset

Dataset Summary

The Mostafa Mahmoud Arabic Speech Dataset is a large-scale Arabic speech corpus containing approximately 187 hours of speech recordings and corresponding transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud.

The dataset was created to support Arabic speech technology research and development, including:

  • —Automatic Speech Recognition (ASR)
  • —Text-to-Speech (TTS)
  • —Speech Foundation Models
  • —Audio-Text Alignment
  • —Speaker Adaptation
  • —Arabic Language Technology Research

Transcriptions were generated and curated by the dataset creator using an AI-assisted transcription pipeline and additional quality-control procedures.

With 187 hours of speech from a single speaker, the dataset is particularly suitable for:

  • —High-quality Arabic TTS voice cloning
  • —Speaker adaptation and speaker representation learning
  • —Long-form ASR training
  • —Foundation model pretraining and fine-tuning
  • —Research on Arabic speech and narration styles

Dataset Statistics

MetricValue
LanguageArabic
SpeakerDr. Mostafa Mahmoud
Total Audio Duration~187 Hours
Number of Speakers1
TasksASR, TTS
Transcript GenerationAI-assisted, creator-curated
FormatAudio + Text

يحتفظ المنشئون الأصليون والقنوات المالكة بكافة الحقوق، ولا يتم ادعاء أي ملكية للملفات الصوتية الأصلية أو المحتوى الصوتي الأساسي.