datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
darija-asr-corpus
Darija ASR Corpus (dataset-core)
Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
This repo contains four source subsets: DODa, DVoice, Wiki, and
YouTube. Each subset carries its own upstream license/terms -- see below --
because they are drawn from four different original projects.
Subsets
Config
Rows
Audio bundled?
Upstream license
Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.MedQA-Darija-MultiLingual
MedQA-Darija-MultiLingual
The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija.
A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region.
Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.doda-darija-cosyvoice2
Dataset Card for DODa Moroccan Darija (CosyVoice2 Ready-to-Train)
Dataset Summary
DODa Moroccan Darija (CosyVoice2 Edition) is a curated, standardized, and tokenized speech dataset engineered specifically for fine-tuning CosyVoice2 on Moroccan Arabic (Darija).
While raw audio datasets typically require extensive preprocessing (sample rate normalization, voice activity detection, multi-speaker segmentation, semantic tokenization, speaker embedding extraction, and… See the full description on the dataset page: https://huggingface.co/datasets/Jip7e/doda-darija-cosyvoice2.Darija3-denoiseddarija_yt_2026
darija_yt_2026
Partition upload generated automatically.
Namespace: ohsn
Repo: ohsn/darija_yt_2026
Video count: 3511
Duration hours: 1565.31
This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline.
darija_speech_to_textdarija_sttdarija-stt-mixThe Darija Speech To Text Dataset is a comprehensive collection designed to support speech recognition tasks for the Darija dialect, it includes audio data totaling 8.23 GB and consists of 13,178 rows of transcribed speech.
This dataset covers a variety of dialects, primarily focusing on Algerian and Moroccan Darija, and also includes slang from other Arabic-speaking countries.
The data has been meticulously gathered from diverse resources to ensure a rich and varied representation of spoken… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/darija-stt-mix.darija-merged-asrdarija-asr-3h
Moroccan Darija ASR — 3 hours
YouTube Moroccan Darija, segmented and filtered, labeled with Gemini 2.5 Pro.
split
hours
clips
train
3.00
1778
validation
0.15
91
silver
0.35
184
Splits are channel-disjoint: silver channels do not appear in train. A same-size random split leaks 100% of silver channels into train.
Columns
id, audio (16 kHz), text (Gemini 2.5 Pro)
channel (YouTube handle)
duration, pesq_hyp (SQUIM, no-reference), num_speakers… See the full description on the dataset page: https://huggingface.co/datasets/01Yassine/darija-asr-3h.DarijaTTS-cleantts_darija
language:
- ar
license: cc-by-4.0
task_categories:
- automatic-speech-recognition
task_ids:
- automatic-speech-recognition
pretty_name: Darija Arabic Speech Dataset
size_categories:
- 1K<n<10K
tags:
- darija
- moroccan-arabic
- arabic
- speech
- asr
- automatic-speech-recognition
- whisper
- morocco
Moroccan Darija Speech Dataset
A speech dataset for Moroccan Arabic (Darija) automatic speech recognition (ASR).
The dataset consists of short audio clips extracted… See the full description on the dataset page: https://huggingface.co/datasets/anassdabaghi/tts_darija.DVOICEv2.0-DarijaDVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/DVOICEv2.0-Darija.darija-tts-8400
Darija TTS 8400
Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV.
All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings.
Write-up of how this data was used: Training a Voice.
At a glance
Clips / hours
8,400 / 20.73
Unique texts
4,800
Voice
Kore (1 speaker)
Sample rate
24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.darija-speech-to-text
Speech To Text Darija dataset
Reupload of adiren7/darija_speech_to_text
darija-asr-benchmark-6speaker
Darija ASR 6-Speaker Benchmark
A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3
female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus),
used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi)
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
Consent and anonymization
Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.darija-stt-dataset
Dataset Card for "darija-stt-dataset"
More Information needed
darija_stt_mixdarija_to_french_speech_to_textMoroccan-Darija-Wiki-Audio-Dataset
Moroccan Darija Wiki Audio Dataset
Overview
The Moroccan Darija Wiki Audio Dataset consists of 551 parallel text and speech samples of Moroccan Darija sourced from Wikipedia Darija . This dataset is designed to support speech recognition, language modeling, and various NLP tasks for Moroccan Darija.
Dataset Source
The data was scraped from Wikipedia (ary) using the WikiScraper tool.
Data Preprocessing
To ensure data quality, we applied… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/Moroccan-Darija-Wiki-Audio-Dataset.DATASET-darija
Darija ASR Dataset
Dataset de reconnaissance automatique de la parole en Darija marocaine.
Description
Langue: Darija marocaine (ary)
Tache: automatic speech recognition
Audio: WAV mono 16 kHz stocke en Parquet
Colonnes: audio, sentence
Structure
Colonne
Type
Description
audio
Audio
Segment audio WAV mono 16 kHz
sentence
string
Transcription en darija
License
CC BY 4.0
darija-asr-cleandarija-dz-tts-v1moroccan-darija-asr-dataset-splitSegmented-Moroccan-Darija-Wiki-Audio-Dataset
Dataset Card for Segmented Moroccan Darija Wiki Dataset
Dataset Summary
This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon).
Each audio is split into segments of up to 30 seconds to make it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/anaszil/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.darija-synthetic-callsDATASET-darija-ASR-clean
Darija ASR Dataset
Dataset de reconnaissance automatique de la parole en Darija marocaine.
Description
Langue: Darija marocaine (ary)
Tache: automatic speech recognition
Audio: WAV mono 16 kHz stocke en Parquet
Colonnes: audio, sentence
Structure
Colonne
Type
Description
audio
Audio
Segment audio WAV mono 16 kHz
sentence
string
Transcription en darija
License
CC BY 4.0
DVOICEv1.1-DarijaDialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/DVOICEv1.1-Darija.darija-tts
Moroccan Darija TTS Dataset
This dataset contains Moroccan Darija speech recordings and their corresponding transcriptions, intended for fine-tuning text-to-speech models.
Dataset Structure
wavs_16k/: Directory containing 16kHz mono 16-bit WAV audio files.
metadata_train.csv: CSV file with training data.
metadata_val.csv: CSV file with validation data.
Usage
Load the dataset using:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Fahd1199/darija-tts.darija-HF-dataset
