datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-audio-collection-moroccan-ameed
Ameed Moroccan Arabic Speech Dataset
Dataset Summary
The Ameed Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 176 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-ameed.arabic-audio-collection-moroccan-wak3i
Mak3i Moroccan Arabic Speech Dataset
Dataset Summary
The Mak3i Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 70 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-wak3i.arabic-audio-collection-moroccan-noone-stories
Noone Stories Moroccan Arabic Speech Dataset
Dataset Summary
The Noone Stories Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 156 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-noone-stories.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.moroccan_amazigh_asrSegmented-Moroccan-Darija-Wiki-Audio-Dataset
Dataset Card for Segmented Moroccan Darija Wiki Dataset
Dataset Summary
This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon).
Each audio is split into segments of up to 30 seconds to make it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/anaszil/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.Segmented-Moroccan-Darija-Wiki-Audio-Dataset
Dataset Card for Segmented Moroccan Darija Wiki Dataset
Dataset Summary
This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon).
Each audio is split into segments of up to 30 seconds to make it suitable… See the full description on the dataset page: https://huggingface.co/datasets/H20-sys/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.
