datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-audio-collection-moroccan-ameed
Ameed Moroccan Arabic Speech Dataset
Dataset Summary
The Ameed Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 176 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-ameed.arabic-audio-collection-moroccan-wak3i
Mak3i Moroccan Arabic Speech Dataset
Dataset Summary
The Mak3i Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 70 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-wak3i.arabic-audio-collection-moroccan-noone-stories
Noone Stories Moroccan Arabic Speech Dataset
Dataset Summary
The Noone Stories Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 156 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-noone-stories.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.Moroccan-Darija-Wiki-Audio-Dataset
Moroccan Darija Wiki Audio Dataset
Overview
The Moroccan Darija Wiki Audio Dataset consists of 551 parallel text and speech samples of Moroccan Darija sourced from Wikipedia Darija . This dataset is designed to support speech recognition, language modeling, and various NLP tasks for Moroccan Darija.
Dataset Source
The data was scraped from Wikipedia (ary) using the WikiScraper tool.
Data Preprocessing
To ensure data quality, we applied… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/Moroccan-Darija-Wiki-Audio-Dataset.moroccan-darija-asr-dataset-splitmoroccan_amazigh_asrSegmented-Moroccan-Darija-Wiki-Audio-Dataset
Dataset Card for Segmented Moroccan Darija Wiki Dataset
Dataset Summary
This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon).
Each audio is split into segments of up to 30 seconds to make it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/anaszil/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.Moroccan-Darija-ASRSegmented-Moroccan-Darija-Wiki-Audio-Dataset
Dataset Card for Segmented Moroccan Darija Wiki Dataset
Dataset Summary
This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon).
Each audio is split into segments of up to 30 seconds to make it suitable… See the full description on the dataset page: https://huggingface.co/datasets/H20-sys/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.Moroccan_Darija_XTTS_MasterMoroccan-STT-Eval-DatasetMoroccan-Darija-HSIOmoroccan_darija_audio
