datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DarijaDz
DarijaDZ
DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content.
Dataset Description
Motivation
Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.Algerian-Darija
Overview
This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs.
The train split consists more then 2k rows of uncleaned text data.
The v1 split consists more than 170k rows of split and partially cleaned text.
Sources
The text data was gathered from:
Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija.
Web Scraping: Content… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/Algerian-Darija.TinyStories-Algerian-Darija
TinyStories Algerian Darija
Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous).
The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.facebook_darija_dataset
Darija Facebook Posts Dataset
Dataset Details
Dataset Description
This dataset consists of more than 5k public posts from Facebook. Each post contains text content, metadata.
This dataset containt more than 400K darija tokens.
Curated by: @abdeljalilELmajjodi
Language(s) (NLP): Multiple (primarily Moroccon Arabic)
Uses
This dataset could be used for:
Training and testing language models on social media content
Analyzing social media posting… See the full description on the dataset page: https://huggingface.co/datasets/abdeljalilELmajjodi/facebook_darija_dataset.Darija-SFT-Mixture
Dataset Card for Darija-SFT-Mixture
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
Darija-SFT-Mixture is a dataset consisting of 458K instruction samples, by consolidating existing Darija language resources, creating novel datasets both manually and synthetically, and translating English instructions under strict quality control.… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/Darija-SFT-Mixture.chat-darija-therapy
Moroccan Darija Therapy Conversations Dataset
This dataset is entirely synthetic and contains no real patient information.
It is provided strictly for research, educational, and experimental purposes and must not be used for clinical, medical, diagnostic, or psychological decision-making.
Citation
If you use this dataset in your research, please cite:
@dataset{moroccan_darija_therapy_conversations,
title={Moroccan Darija Therapy Conversations},
author={Jamal… See the full description on the dataset page: https://huggingface.co/datasets/yibba/chat-darija-therapy.Darija-Stories-Dataset
Dataset Card for "Darija-Stories-Dataset"
Darija (Moroccan Arabic) Stories Dataset is a large-scale collection of stories written in Moroccan Arabic dialect (Darija).
Dataset Description
Darija (Moroccan Arabic) Stories Dataset contains a diverse range of stories that provide insights into Moroccan culture, traditions, and everyday life. The dataset consists of textual content from various chapters, including narratives, dialogues, and descriptions. Each story chapter is… See the full description on the dataset page: https://huggingface.co/datasets/alielfilali01/Darija-Stories-Dataset.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.algerian-darija-dictionary-v1
Description
This dataset is a comprehensive dictionary of Algerian Darija words and popular proverbs. It aims to document the rich linguistic heritage and daily expressions used in Algeria.
⚠️ Quality Warning
Please note that the current version of the dataset contains noise and spam. It is a work in progress and is not yet 100% accurate.
Dataset Statistics
Number of entries: 4,636
Columns: word, french_writing, definition, example… See the full description on the dataset page: https://huggingface.co/datasets/awras/algerian-darija-dictionary-v1.Moroccan-Darija-Instruct-573K
Moroccan Darija Instruct 573K
Created by Lyte
If you use this dataset, please credit: Lyte/Moroccan-Darija-Instruct-573K
A synthetic instruction-tuning dataset of 573,175 question-answer pairs written entirely in Moroccan Darija (الدّارجة المغربية).
Dataset Summary
Metric
Value
Total rows
573,175
Original pairs
300,194
Augmented variants
272,981
Total words
~19.9M
Est. tokens
~13.3M
File size
399 MB
MSA contamination
0%
Exact duplicates
0%… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/Moroccan-Darija-Instruct-573K.Moroccan-Darija-QA
Moroccan Darija Q&A Dataset
A comprehensive question-answer dataset in Moroccan Darija (Moroccan Arabic dialect) covering various topics of daily life, culture, and practical knowledge.
📊 Dataset Overview
This dataset contains 3,470 question-answer pairs in Moroccan Darija organized across 3 configurations:
🔗 Default: 2,026 standard Q&A pairs
🌍 Translated: 1,300 translated content pairs
🧠 Reasoning: 144 reasoning-based Q&A with thinking process
Topics… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/Moroccan-Darija-QA.moroccan-darija-dataset
Moroccan Darija Comments Dataset 🇲🇦
A large-scale dataset of 130,851 Darija (Moroccan Arabic) comments collected from YouTube and synthetic generation, designed for NLP tasks on the Moroccan dialect.
Dataset Description
Overview
This dataset contains authentic and augmented Darija comments covering diverse topics including sports, music, cuisine, news, culture, humor, and daily life in Morocco. It is built for training and evaluating NLP models on Moroccan… See the full description on the dataset page: https://huggingface.co/datasets/IlyasFardaouixx/moroccan-darija-dataset.darija-combined-dataset
darija-combined-dataset
Dataset Description
This dataset contains Darija (Moroccan Arabic) text data collected from multiple sources and processed for machine learning tasks.
Dataset Summary
Total texts: 35,562
Total characters: 5,891,009
Average text length: 165.7 characters
Number of unique datasets: 8
Created: 2025-07-28T15:03:41.001214
Processing time: 18.93 seconds
Source Distribution
Darija Pattern Dataset: 1,076 texts
Audio Transcriptions:… See the full description on the dataset page: https://huggingface.co/datasets/mohamed-stifi/darija-combined-dataset.Algerian-Darija
Overview
This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs.
The train split consists more then 2k rows of uncleaned text data.
The v1 split consists more than 170k rows of split and partially cleaned text.
Sources
The text data was gathered from:
Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija.
Web Scraping: Content… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/Algerian-Darija.DarijaAlpacaEval
Dataset Card for DarijaAlpacaEval
Dataset Summary
DarijaAlpacaEval is an evaluation dataset designed to assess the performance of large language models on instruction-following tasks in Moroccan Darija, a variety of Arabic. It is adapted from the AlpacaEval dataset and consists of instructions provided in Moroccan Darija. The dataset aims to provide a culturally relevant benchmark for evaluating language models' capabilities in instructions following and responses… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaAlpacaEval.DoDA_sentences_darija_englishDataset of english/moroccan arabic translations, from DoDA dataset.
Health_QA_Darija
Health QA Darija — Medical QA in Moroccan Arabic (الدارجة المغربية)
Dataset Description
A curated dataset of 8,129 medical question-answer pairs in Moroccan Darija (الدارجة المغربية). Each entry contains a patient scenario, a focused medical question, and a doctor's response — all in authentic Darija. Enriched with named medical entities (symptoms, diseases, medications, tests).
🇲🇦 First large-scale medical QA dataset in Moroccan Darija — addressing the critical gap in… See the full description on the dataset page: https://huggingface.co/datasets/Kakyoin03/Health_QA_Darija.DarijaStory
Dataset Card for DarijaStory
Dataset Summary
DarijaStory is a story completion dataset. It consists of 4,392 long stories scraped from 9esa, a website featuring a variety of stories written in Moroccan Darija.
Supported Tasks and Leaderboards
Task Category: Conditional Text Generation
Task: Story Completion in Moroccan Darija
Languages
The dataset is available in Moroccan Arabic (Darija).
Dataset Structure
Data Instances
Each… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaStory.morocco_tourism_darija
🇲🇦 Moroccan Tourism Darija Dialogues (maroc_tourism_darja)
Repository | datasets/maroc_tourism_darjaSize | 1 000 dialogues (≈ 14 k turns)Language | Darija (Moroccan Arabic)Topics | Transport · Accommodation · Culture · GastronomyLicense | CC‑BY‑4.0Format | JSONL (one dialogue per line)
1. Dataset Summary
maroc_tourism_darja is a synthetic, high‑quality collection of tourist‑guide conversations in Moroccan Darija.Each dialogue is a short two‑turn exchange in which a… See the full description on the dataset page: https://huggingface.co/datasets/oabai/morocco_tourism_darija.moroccan-darija-therapy-conversations
Moroccan Darija Therapy Conversations Dataset
Description
This dataset contains approximately 15,000 synthetic therapy-style conversations
between a patient / User and a therapist / Assistant, written in Darija (Moroccan Arabic).
All conversations were generated using a Large Language Model (LLM) and are intended
to support research and development in conversational AI, low-resource NLP, and
Arabic dialect modeling.
Important DisclaimerThis dataset is entirely synthetic… See the full description on the dataset page: https://huggingface.co/datasets/yibba/moroccan-darija-therapy-conversations.fw-darija
Gherbal’ing Multilingual Fineweb 2 🍵
Following up on their previous release, the fineweb team has been hard at work on their upcoming multilingual fineweb dataset which contains a massive collection of 50M+ sentences across 100+ languages. The data, sourced from the Common Crawl corpus, has been classified into these languages using GlotLID, a model able to recognize more than 2000 languages. The performance of GlotLID is quite impressive, considering the complexity of the task. It… See the full description on the dataset page: https://huggingface.co/datasets/sawalni-ai/fw-darija.darija_bible
Darija Bible — Moroccan Standard Translation (MSTD)
The full New Testament translated into Moroccan Darija
(الترجمة المغربية القياسية, MSTD).
Moroccan Darija is a low-resource spoken Arabic variety. This dataset is published
here as a research artifact for language modeling, fine-tuning, machine translation
between Darija and other languages, and evaluation of multilingual / Arabic-dialect
models.
⚠️ Copyright notice — please read before using
This dataset reproduces… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/darija_bible.whisper-darija
