datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
100-hour-Egyptian-dataset-single-speaker
Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus
A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts.
Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data.
Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home.
This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.egyptian-arabic-speechEgyptian-ASR-MGB-3
Egyptian Arabic dialect automatic speech recognition
Dataset Summary
This dataset was collected, cleaned and adjusted for huggingface hub and ready to be used for whisper finetunning/training.
From MGB-3 website:
The MGB-3 is using 16 hours multi-genre data collected from different YouTube channels. The 16 hours have been manually transcribed.
The chosen Arabic dialect for this year is Egyptian.
Given that dialectal Arabic has no orthographic rules, each program has… See the full description on the dataset page: https://huggingface.co/datasets/MightyStudent/Egyptian-ASR-MGB-3.Egyptian-Arabic-Lectures
Egyptian Arabic Lectures Dataset
The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts.
Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.egyptian-arabic-400kEgyptian_Dialect
Egyptian Arabic Speech Dataset
Dataset Description
This dataset contains 2,438 short audio segments in Egyptian Arabic paired with transcriptions.
The dataset was created for fine-tuning Automatic Speech Recognition (ASR) models on conversational Egyptian Arabic.
Each example contains:
WAV audio
Egyptian Arabic transcription
Dataset Creation
Source
The audio was collected from publicly available YouTube videos featuring native… See the full description on the dataset page: https://huggingface.co/datasets/Kyrillos2001/Egyptian_Dialect.dahab-egyptian-female-tts
Dahab — Egyptian Arabic, single female speaker
134.7 hours across 59,505 clips of Egyptian (Cairene) Arabic from one
female speaker, at 24 kHz mono. 26,741 clips (44.9%) carry diacritized
transcripts. Built for TTS fine-tuning.
Segmented from a single YouTube cooking channel, so the register is
conversational instructional speech throughout.
Structure
The train split is stored in self-contained Parquet shards. Each row contains an audio object with embedded WAV… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/dahab-egyptian-female-tts.egyptian-arabic-youtube-ttsEgyptian-ASR-MGB-3
Egyptian Arabic dialect automatic speech recognition
Dataset Summary
This dataset was collected, cleaned and adjusted for huggingface hub and ready to be used for whisper finetunning/training.
From MGB-3 website:
The MGB-3 is using 16 hours multi-genre data collected from different YouTube channels. The 16 hours have been manually transcribed.
The chosen Arabic dialect for this year is Egyptian.
Given that dialectal Arabic has no orthographic rules, each program has… See the full description on the dataset page: https://huggingface.co/datasets/severo/Egyptian-ASR-MGB-3.Egyptian_TTS3RSmoustafa-sadek-egyptian-tts
Moustafa-Sadek Egyptian-Arabic TTS
2883 clips, 24 kHz mono WAV in audio/, transcripts (Deepgram nova-3, tashkeel-stripped, lexicon-normalized) in transcripts.jsonl.
Each line: {id, audio, text, dur, language_id}. Derived from a public YouTube channel — credit the creator; non-commercial.
masri-podcast-300h-egyptian-tts
Masri Podcast - Egyptian Arabic TTS corpus
Multi-speaker Egyptian Arabic speech built from public Egyptian podcast episodes,
cut into single-speaker clips with word-level timings.
Format
24 kHz mono 16-bit FLAC, EBU R128 loudness-normalised, DeepFilterNet3 denoised
clips 3-15 s, cut on pause boundaries, never mid-word
word-level start/end for every word
speaker_id is global across episodes (ECAPA embeddings + agglomerative clustering)
Splits… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/masri-podcast-300h-egyptian-tts.Egyptian-dialect-8k
Egyptian Dialect Speech Dataset (8k)
Dataset Description
This dataset contains around 8,000 cleaned speech samples in the Egyptian Arabic dialect. It is optimized for speech recognition (ASR), text-to-speech (TTS), and audio processing tasks tailored specifically for Egyptian dialect applications.
Primary Language: Egyptian Arabic (ar-EG)
Total Samples: ~8,000 audio clips with corresponding text transcriptions.
Source: Cleaned and processed from… See the full description on the dataset page: https://huggingface.co/datasets/ismailelsayedeltanja/Egyptian-dialect-8k.egyptian-speech-ttsEgyptian-dialect-100kEgyptian-Speech-Clean-MGB3
🏛️ Dataset Card for MGB3-Egyptian-Clean
Dataset Summary
This dataset is a refined and enhanced version of the MGB-3 (Multi-Genre Broadcast) corpus, specifically focused on the Egyptian Arabic dialect. It has been meticulously preprocessed to be "TTS-ready" by combining advanced deep-learning denoising with custom linguistic text normalization.
🛠️ Preprocessing Pipeline
To ensure the highest quality for generative speech tasks (like VITS or MMS finetuning)… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/Egyptian-Speech-Clean-MGB3.egyptian_arabic_speech_zaidegyptianDatasetegyptian-voice-dataset-finalEgyptian_dialect-filtered-dataEgyptian-Arabic-synthetic-sttsynthetic_egyptian-arabic-ttsEgyptian-Speech-Audio-Text-1kEgyptian-TTS-Evaluation-EGTTSegyptian-voice-dataset-2egyptian-voice-datasetegyptian-voice-dataset-final2egyptian_ttsegyptian_myvoice-last_testegyptian_voice_lev
