datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
darija-asr-corpus
Darija ASR Corpus (dataset-core)
Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
This repo contains four source subsets: DODa, DVoice, Wiki, and
YouTube. Each subset carries its own upstream license/terms -- see below --
because they are drawn from four different original projects.
Subsets
Config
Rows
Audio bundled?
Upstream license
Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.DarijaDz
DarijaDZ
DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content.
Dataset Description
Motivation
Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.DarijaDZ-DialectID
DarijaDZ Dialect Identification
DarijaDZ-DialectID is a labeled dataset for classifying Algerian
online text into one of six dialect/language classes: darija, msa,
arabize, french, english, code_switch. It is part of
DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija.
Dataset Description
Motivation
Algeria's online text is a mix of several dialects and scripts --
Algerian Darija (Arabic script), Modern Standard Arabic, Arabizi… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDZ-DialectID.darija-tts-8400
Darija TTS 8400
Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV.
All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings.
Write-up of how this data was used: Training a Voice.
At a glance
Clips / hours
8,400 / 20.73
Unique texts
4,800
Voice
Kore (1 speaker)
Sample rate
24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.darija-asr-benchmark-6speaker
Darija ASR 6-Speaker Benchmark
A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3
female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus),
used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi)
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
Consent and anonymization
Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.chat-darija-therapy
Moroccan Darija Therapy Conversations Dataset
This dataset is entirely synthetic and contains no real patient information.
It is provided strictly for research, educational, and experimental purposes and must not be used for clinical, medical, diagnostic, or psychological decision-making.
Citation
If you use this dataset in your research, please cite:
@dataset{moroccan_darija_therapy_conversations,
title={Moroccan Darija Therapy Conversations},
author={Jamal… See the full description on the dataset page: https://huggingface.co/datasets/yibba/chat-darija-therapy.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-moroccan_darija_prompts & trilingual_codeswitch_chat (augmented)
This dataset consists of short conversational prompts written in Moroccan Darija, covering topics like shopping, social interactions, and daily inquiries. Each entry contains a single prompt with a null completion, indicating it is likely intended for instruction tuning or completion generation tasks. The content… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented.moroccan-darija-llm-datasetdarija-therapy-qa
Moroccan Darija Therapy Conversations Dataset
This dataset is entirely synthetic and contains no real patient information.
It is provided strictly for research, educational, and experimental purposes and must not be used for clinical, medical, diagnostic, or psychological decision-making.
Citation
If you use this dataset in your research, please cite:
@dataset{moroccan_darija_therapy_conversations,
title={Moroccan Darija Therapy Conversations},
author={Jamal… See the full description on the dataset page: https://huggingface.co/datasets/yibba/darija-therapy-qa.learning-darijaEvaluatif-DarijaMinddarija_finetune_data_for_gpt4ominidarija-therapy-qamo_darija_mergednorhten_darija_dialect_data_30Kdarija
