moroccan-arabic
arabic-audio-collection-moroccan-ameed
Ameed Moroccan Arabic Speech Dataset
Dataset Summary
The Ameed Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 176 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-ameed.arabic-audio-collection-moroccan-wak3i
Mak3i Moroccan Arabic Speech Dataset
Dataset Summary
The Mak3i Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 70 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-wak3i.FineWeb2-Moroccan-Arabic
FineWeb2 Moroccan Arabic
🇲🇦 This is the Moroccan Arabic Portion of The FineWeb2 Dataset.
🇲🇦 The Moroccan Arabic language, represented by the ISO 639-3 code ary, is a member of the Afro-Asiatic language family and utilizes the Arabic script.
🇲🇦 Known within subsets as ary_Arab, this language boasts an extensive corpus of over 1.7 billion words across more than 6.1 million documents, collectively occupying a disk size of approximately 5.79 GB.
Purpose of This Repository… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-Moroccan-Arabic.arabic-audio-collection-moroccan-noone-stories
Noone Stories Moroccan Arabic Speech Dataset
Dataset Summary
The Noone Stories Moroccan Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 156 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-moroccan-noone-stories.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.Moroccan_Arabic_Wikipedia_20230101_nobots
Dataset Card for "Moroccan_Arabic_Wikipedia_20230101_nobots"
This dataset is created using the Moroccan Arabic Wikipedia articles (after removing bot-generated articles), downloaded on the 1st of January 2023, processed using Gensim Python library, and preprocessed using tr Linux/Unix utility and CAMeLTools Python toolkit for Arabic NLP. This dataset was used to train this Moroccan Arabic Wikipedia Masked Language Model: SaiedAlshahrani/arywiki_20230101_roberta_mlm_nobots.
For more… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/Moroccan_Arabic_Wikipedia_20230101_nobots.
