datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
darija-asr-corpus
Darija ASR Corpus (dataset-core)
Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
This repo contains four source subsets: DODa, DVoice, Wiki, and
YouTube. Each subset carries its own upstream license/terms -- see below --
because they are drawn from four different original projects.
Subsets
Config
Rows
Audio bundled?
Upstream license
Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.MedQA-Darija-MultiLingual
MedQA-Darija-MultiLingual
The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija.
A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region.
Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.darija_yt_2026
darija_yt_2026
Partition upload generated automatically.
Namespace: ohsn
Repo: ohsn/darija_yt_2026
Video count: 3511
Duration hours: 1565.31
This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline.
darija_speech_to_textdarija-stt-mixThe Darija Speech To Text Dataset is a comprehensive collection designed to support speech recognition tasks for the Darija dialect, it includes audio data totaling 8.23 GB and consists of 13,178 rows of transcribed speech.
This dataset covers a variety of dialects, primarily focusing on Algerian and Moroccan Darija, and also includes slang from other Arabic-speaking countries.
The data has been meticulously gathered from diverse resources to ensure a rich and varied representation of spoken… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/darija-stt-mix.darija-asr-3h
Moroccan Darija ASR — 3 hours
YouTube Moroccan Darija, segmented and filtered, labeled with Gemini 2.5 Pro.
split
hours
clips
train
3.00
1778
validation
0.15
91
silver
0.35
184
Splits are channel-disjoint: silver channels do not appear in train. A same-size random split leaks 100% of silver channels into train.
Columns
id, audio (16 kHz), text (Gemini 2.5 Pro)
channel (YouTube handle)
duration, pesq_hyp (SQUIM, no-reference), num_speakers… See the full description on the dataset page: https://huggingface.co/datasets/01Yassine/darija-asr-3h.darija-speech-to-text
Speech To Text Darija dataset
Reupload of adiren7/darija_speech_to_text
darija-asr-benchmark-6speaker
Darija ASR 6-Speaker Benchmark
A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3
female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus),
used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi)
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
Consent and anonymization
Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.DATASET-darija
Darija ASR Dataset
Dataset de reconnaissance automatique de la parole en Darija marocaine.
Description
Langue: Darija marocaine (ary)
Tache: automatic speech recognition
Audio: WAV mono 16 kHz stocke en Parquet
Colonnes: audio, sentence
Structure
Colonne
Type
Description
audio
Audio
Segment audio WAV mono 16 kHz
sentence
string
Transcription en darija
License
CC BY 4.0
Segmented-Moroccan-Darija-Wiki-Audio-Dataset
Dataset Card for Segmented Moroccan Darija Wiki Dataset
Dataset Summary
This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon).
Each audio is split into segments of up to 30 seconds to make it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/anaszil/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.DATASET-darija-ASR-clean
Darija ASR Dataset
Dataset de reconnaissance automatique de la parole en Darija marocaine.
Description
Langue: Darija marocaine (ary)
Tache: automatic speech recognition
Audio: WAV mono 16 kHz stocke en Parquet
Colonnes: audio, sentence
Structure
Colonne
Type
Description
audio
Audio
Segment audio WAV mono 16 kHz
sentence
string
Transcription en darija
License
CC BY 4.0
wikitongues-darija
Wikitongues-Darija
This is a small test dataset for Automatic Speech Recognition in Darija language, built from 2 captioned videos of the WikiTongues project:
nawal
anass
Process:
each webm video has been converted to monochannel 16khz wav files with ffmpeg :
ffmpeg -i WIKITONGUES-_Nawal_speaking_Moroccan_Arabic.webm.1080p.vp9.webm -ar 16000 -ac 1 nawal.wav
each audio has been cut in samples of less than 30 seconds audio according to the captions timestamps. The script may be… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/wikitongues-darija.darija-youtube-dataset
Darija YouTube Dataset
A dataset of Moroccan Arabic (Darija) speech scraped from YouTube, with transcriptions in Arabic script, Latin script (Arabizi), and English translations.
Dataset Description
This dataset contains sentence-level audio segments of Darija speech, paired with:
Arabic transcription (modern Moroccan Arabic script)
Latin transliteration (Arabizi format: 3=ع, 7=ح, 9=ق, etc.)
English translation
Columns
Column
Type
Description
audio… See the full description on the dataset page: https://huggingface.co/datasets/DrIAmed/darija-youtube-dataset.Segmented-Moroccan-Darija-Wiki-Audio-Dataset
Dataset Card for Segmented Moroccan Darija Wiki Dataset
Dataset Summary
This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon).
Each audio is split into segments of up to 30 seconds to make it suitable… See the full description on the dataset page: https://huggingface.co/datasets/H20-sys/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.
