CoolFace
Datasetpublic

oddadmix/arabic-audio-collection-sudanese-ahmed-gobara

Ahmed Gobara Sudanese Arabic Speech Dataset Dataset Summary The Ahmed Gobara Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 19 hours of speech recordings and corresponding transcripts. While compact, the dataset offers a clean, consistent single-speaker resource in Sudanese Arabic — an Arabic variety with very few open speech resources — making it especially valuable for voice cloning, speaker adaptation… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-ahmed-gobara.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
1likes122downloads
Dataset Card

Ahmed Gobara Sudanese Arabic Speech Dataset

Dataset Summary

The Ahmed Gobara Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 19 hours of speech recordings and corresponding transcripts.

While compact, the dataset offers a clean, consistent single-speaker resource in Sudanese Arabic — an Arabic variety with very few open speech resources — making it especially valuable for voice cloning, speaker adaptation, and dialect-focused fine-tuning experiments.

Alongside the spoken Arabic text, the transcripts capture paralinguistic vocalizations, breathing, and emotional cues, enabling the development of expressive, natural-sounding AI models that go beyond flat dictation.

The dataset was created to support Arabic speech technology research and development, including:

  • —Automatic Speech Recognition (ASR)
  • —Text-to-Speech (TTS) and Expressive TTS
  • —Speech Foundation Models
  • —Audio-Text Alignment
  • —Speaker Adaptation
  • —Paralinguistic and Emotion Recognition
  • —Sudanese Dialect and Arabic Language Technology Research

Transcriptions were generated and curated by the dataset creator using an AI-assisted transcription pipeline and additional quality-control procedures to ensure the accurate logging of non-verbal tokens.

Non-Verbal Vocalization Tags

The transcripts include a comprehensive set of non-verbal tokens to capture the true nuance of human speech and breathing. The supported tags are:

<laugh> <cry> <weep> <sob> <scream> <shout> <whisper> <sigh> <gasp> <groan> <moan> <pause> <hes> <stutter> <breath> <sniff> <cough> <throat_clear>

Research & Application Use Cases

With ~19 hours of single-speaker Sudanese Arabic speech and granular non-verbal tagging, the dataset is particularly suitable for:

  • —Sudanese Arabic TTS and voice cloning: A consistent single-speaker voice for dialect-aware, expressive synthesis.
  • —Speaker adaptation: Fine-tuning existing Arabic ASR/TTS models to a Sudanese speaker with limited data.
  • —ASR fine-tuning and evaluation for Sudanese Arabic.
  • —Foundation model fine-tuning.
  • —Research on Sudanese Arabic speech, narration styles, and paralinguistics.

Dataset Statistics

MetricValue
LanguageArabic (Sudanese dialect)
SpeakerAhmed Gobara
Total Audio Duration~19 Hours
Number of Speakers1
TasksASR, TTS, Expressive TTS, Paralinguistics
Transcript GenerationAI-assisted, creator-curated
Special FeaturesRich non-verbal and emotional transcription tags
FormatAudio + Text

يحتفظ المنشئون الأصليون والقنوات المالكة بكافة الحقوق، ولا يتم ادعاء أي ملكية للملفات الصوتية الأصلية أو المحتوى الصوتي الأساسي.