oddadmix/arabic-audio-collection-syrian-podcast
Syrian Postcast Arabic Speech Dataset Dataset Summary The Syrian Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 116 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-syrian-podcast.
Syrian Postcast Arabic Speech Dataset
Dataset Summary
The Syrian Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 116 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture paralinguistic vocalizations, breathing, and emotional cues. This makes it uniquely suited for developing highly expressive, natural-sounding AI models that go beyond flat dictation.
The dataset was created to support Arabic speech technology research and development, including:
- Automatic Speech Recognition (ASR)
- Text-to-Speech (TTS) and Expressive TTS
- Speech Foundation Models
- Audio-Text Alignment
- Speaker Adaptation
- Paralinguistic and Emotion Recognition
- Arabic Language Technology Research
Transcriptions were generated and curated by the dataset creator using an AI-assisted transcription pipeline and additional quality-control procedures to ensure the accurate logging of non-verbal tokens.
Non-Verbal Vocalization Tags
The transcripts include a comprehensive set of non-verbal tokens to capture the true nuance of human speech and breathing. The supported tags are:
<laugh> <cry> <weep> <sob> <scream> <shout> <whisper> <sigh> <gasp> <groan> <moan> <pause> <hes> <stutter> <breath> <sniff> <cough> <throat_clear>
Research & Application Use Cases
With 116 hours of speech from a single speaker combined with granular non-verbal tagging, the dataset is particularly suitable for:
- High-fidelity conversational AI: Training models that understand and generate natural human hesitations, pauses, and breaths.
- Highly expressive Arabic TTS: Voice cloning that can synthesize emotion (laughing, whispering, crying) naturally.
- Speaker adaptation and speaker representation learning.
- Long-form ASR training with robust noise/vocalization handling.
- Foundation model pretraining and fine-tuning.
- Research on Arabic speech, narration styles, and paralinguistics.
Dataset Statistics
يحتفظ المنشئون الأصليون والقنوات المالكة بكافة الحقوق، ولا يتم ادعاء أي ملكية للملفات الصوتية الأصلية أو المحتوى الصوتي الأساسي.
