datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DarijaMMLU
Dataset Card for DarijaMMLU
Dataset Summary
DarijaMMLU is an evaluation benchmark designed to assess large language models' (LLM) performance in Moroccan Darija, a variety of Arabic. It consists of 22,027 multiple-choice questions, translated from selected subsets of the Massive Multitask Language Understanding (MMLU) and ArabicMMLU benchmarks to measure model performance on 44 subjects in Darija.
Supported Tasks
Task Category: Multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaMMLU.TinyStories-Algerian-Darijadarija-englishThis work is part of DODa.
darija-asr-corpus
Darija ASR Corpus (dataset-core)
Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
This repo contains four source subsets: DODa, DVoice, Wiki, and
YouTube. Each subset carries its own upstream license/terms -- see below --
because they are drawn from four different original projects.
Subsets
Config
Rows
Audio bundled?
Upstream license
Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.MedQA-Darija-MultiLingual
MedQA-Darija-MultiLingual
The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija.
A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region.
Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.DarijaDz
DarijaDZ
DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content.
Dataset Description
Motivation
Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.doda-darija-cosyvoice2
Dataset Card for DODa Moroccan Darija (CosyVoice2 Ready-to-Train)
Dataset Summary
DODa Moroccan Darija (CosyVoice2 Edition) is a curated, standardized, and tokenized speech dataset engineered specifically for fine-tuning CosyVoice2 on Moroccan Arabic (Darija).
While raw audio datasets typically require extensive preprocessing (sample rate normalization, voice activity detection, multi-speaker segmentation, semantic tokenization, speaker embedding extraction, and… See the full description on the dataset page: https://huggingface.co/datasets/Jip7e/doda-darija-cosyvoice2.Algerian-Darija
Overview
This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs.
The train split consists more then 2k rows of uncleaned text data.
The v1 split consists more than 170k rows of split and partially cleaned text.
Sources
The text data was gathered from:
Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija.
Web Scraping: Content… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/Algerian-Darija.Darija3-denoiseddarija_yt_2026
darija_yt_2026
Partition upload generated automatically.
Namespace: ohsn
Repo: ohsn/darija_yt_2026
Video count: 3511
Duration hours: 1565.31
This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline.
DarijaHellaSwag
Dataset Card for DarijaHellaSwag
Dataset Summary
DarijaHellaSwag is a challenging multiple-choice benchmark designed to evaluate machine reading comprehension and commonsense reasoning in Moroccan Darija. It is a translated version of the HellaSwag validation set, which presents scenarios where models must choose the most plausible continuation of a passage from four options.
Supported Tasks
Task Category: Multiple-choice question answering
Task: Answering… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaHellaSwag.DarijaBench
DarijaBench: A Comprehensive Evaluation Dataset for Summarization, Translation, and Sentiment Analysis in Darija
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
The Moroccan Arabic dialect, commonly referred to as Darija, is a widely spoken but understudied variant of Arabic with distinct linguistic features that differ… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaBench.darija_sttdarija-stt-mixThe Darija Speech To Text Dataset is a comprehensive collection designed to support speech recognition tasks for the Darija dialect, it includes audio data totaling 8.23 GB and consists of 13,178 rows of transcribed speech.
This dataset covers a variety of dialects, primarily focusing on Algerian and Moroccan Darija, and also includes slang from other Arabic-speaking countries.
The data has been meticulously gathered from diverse resources to ensure a rich and varied representation of spoken… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/darija-stt-mix.TinyStories-Algerian-Darija
TinyStories Algerian Darija
Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous).
The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.darija_speech_to_textdarija-merged-asrdarija-asr-3h
Moroccan Darija ASR — 3 hours
YouTube Moroccan Darija, segmented and filtered, labeled with Gemini 2.5 Pro.
split
hours
clips
train
3.00
1778
validation
0.15
91
silver
0.35
184
Splits are channel-disjoint: silver channels do not appear in train. A same-size random split leaks 100% of silver channels into train.
Columns
id, audio (16 kHz), text (Gemini 2.5 Pro)
channel (YouTube handle)
duration, pesq_hyp (SQUIM, no-reference), num_speakers… See the full description on the dataset page: https://huggingface.co/datasets/01Yassine/darija-asr-3h.moroccan-darija-youtube-subtitles
Moroccan Darija YouTube Subtitles Dataset
This dataset contains subtitles from YouTube videos in Moroccan Darija, a colloquial Arabic dialect spoken in Morocco. The subtitles were collected from several popular Moroccan YouTube channels, providing a diverse set of transcriptions in the Darija language.
Dataset Description
The dataset is provided as a CSV file, where each row represents a YouTube video and contains the following columns:
video_id: The unique identifier of… See the full description on the dataset page: https://huggingface.co/datasets/bourbouh/moroccan-darija-youtube-subtitles.DarijaDZ-DialectID
DarijaDZ Dialect Identification
DarijaDZ-DialectID is a labeled dataset for classifying Algerian
online text into one of six dialect/language classes: darija, msa,
arabize, french, english, code_switch. It is part of
DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija.
Dataset Description
Motivation
Algeria's online text is a mix of several dialects and scripts --
Algerian Darija (Arabic script), Modern Standard Arabic, Arabizi… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDZ-DialectID.darija_english
Dataset Card for atlasia/darija-english
Dataset Details
Dataset Description
A compilation of Darija-English pairs curated by AtlasIA.
Curated by: AtlasIA
Language(s) (NLP): Moroccan Darija, English
License: CC-by-NC-4.0
Darija sentences sources (additionally to the web):
doda: AtlasIA platform contributions
stories: Mixed Arabic Datasets
transliteration: AtlasIA x DODa. Can be used for transliteration task.
DVOICEv2.0-DarijaDVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/DVOICEv2.0-Darija.DarijaTTS-v0.2
How to Use the DarijaTTS-v0.2 Dataset
Code:
import IPython.display as ipd
import io
import numpy as np
import tempfile
import wave
import os
from datasets import load_dataset
from IPython.display import Audio
# Load the DarijaTTS-v0.2 dataset with streaming
streaming_dataset = load_dataset("Lyte/DarijaTTS-v0.2", streaming=True)
print("Dataset loaded with streaming:")
print(streaming_dataset)
# Function to play audio from a streaming dataset element
def… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/DarijaTTS-v0.2.DarijaTTS-cleanenglish-to-darija-arabic-script-formattedFineTranslations_Darijadarija-tts-8400
Darija TTS 8400
Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV.
All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings.
Write-up of how this data was used: Training a Voice.
At a glance
Clips / hours
8,400 / 20.73
Unique texts
4,800
Voice
Kore (1 speaker)
Sample rate
24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.tts_darija
language:
- ar
license: cc-by-4.0
task_categories:
- automatic-speech-recognition
task_ids:
- automatic-speech-recognition
pretty_name: Darija Arabic Speech Dataset
size_categories:
- 1K<n<10K
tags:
- darija
- moroccan-arabic
- arabic
- speech
- asr
- automatic-speech-recognition
- whisper
- morocco
Moroccan Darija Speech Dataset
A speech dataset for Moroccan Arabic (Darija) automatic speech recognition (ASR).
The dataset consists of short audio clips extracted… See the full description on the dataset page: https://huggingface.co/datasets/anassdabaghi/tts_darija.darija-speech-to-text
Speech To Text Darija dataset
Reupload of adiren7/darija_speech_to_text
darija_pairs_multilang_dataset
