CoolFace
Datasetpublic

oddadmix/arabic-audio-collection-sudanese-nuuar

Nuuar Sudanese Arabic Speech Dataset Dataset Summary The Nuuar Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 75 hours of speech recordings and corresponding transcripts. Sudanese Arabic remains one of the most underrepresented Arabic varieties in speech technology. This dataset directly addresses that gap by providing long-form, natural, dialectal Sudanese speech from a single consistent speaker, making… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-nuuar.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes127downloads
Dataset Card

Nuuar Sudanese Arabic Speech Dataset

Dataset Summary

The Nuuar Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 75 hours of speech recordings and corresponding transcripts.

Sudanese Arabic remains one of the most underrepresented Arabic varieties in speech technology. This dataset directly addresses that gap by providing long-form, natural, dialectal Sudanese speech from a single consistent speaker, making it well suited for both recognition and synthesis research.

Alongside the spoken Arabic text, the transcripts capture paralinguistic vocalizations, breathing, and emotional cues, enabling the development of expressive, natural-sounding AI models that go beyond flat dictation.

The dataset was created to support Arabic speech technology research and development, including:

  • —Automatic Speech Recognition (ASR)
  • —Text-to-Speech (TTS) and Expressive TTS
  • —Speech Foundation Models
  • —Audio-Text Alignment
  • —Speaker Adaptation
  • —Paralinguistic and Emotion Recognition
  • —Sudanese Dialect and Arabic Language Technology Research

Transcriptions were generated and curated by the dataset creator using an AI-assisted transcription pipeline and additional quality-control procedures to ensure the accurate logging of non-verbal tokens.

Non-Verbal Vocalization Tags

The transcripts include a comprehensive set of non-verbal tokens to capture the true nuance of human speech and breathing. The supported tags are:

<laugh> <cry> <weep> <sob> <scream> <shout> <whisper> <sigh> <gasp> <groan> <moan> <pause> <hes> <stutter> <breath> <sniff> <cough> <throat_clear>

Research & Application Use Cases

With ~75 hours of single-speaker Sudanese Arabic speech and granular non-verbal tagging, the dataset is particularly suitable for:

  • —Sudanese Arabic ASR: Training and fine-tuning recognition models on an underrepresented dialect.
  • —Expressive Sudanese TTS: Voice cloning and dialect-aware synthesis with natural hesitations, pauses, and breaths.
  • —Speaker adaptation and speaker representation learning.
  • —Long-form ASR training with robust noise/vocalization handling.
  • —Foundation model pretraining and fine-tuning.
  • —Research on Sudanese Arabic speech, narration styles, and paralinguistics.

Dataset Statistics

MetricValue
LanguageArabic (Sudanese dialect)
SpeakerNuuar
Total Audio Duration~75 Hours
Number of Speakers1
TasksASR, TTS, Expressive TTS, Paralinguistics
Transcript GenerationAI-assisted, creator-curated
Special FeaturesRich non-verbal and emotional transcription tags
FormatAudio + Text

يحتفظ المنشئون الأصليون والقنوات المالكة بكافة الحقوق، ولا يتم ادعاء أي ملكية للملفات الصوتية الأصلية أو المحتوى الصوتي الأساسي.