datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DarijaMMLU
Dataset Card for DarijaMMLU
Dataset Summary
DarijaMMLU is an evaluation benchmark designed to assess large language models' (LLM) performance in Moroccan Darija, a variety of Arabic. It consists of 22,027 multiple-choice questions, translated from selected subsets of the Massive Multitask Language Understanding (MMLU) and ArabicMMLU benchmarks to measure model performance on 44 subjects in Darija.
Supported Tasks
Task Category: Multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaMMLU.MedQA-Darija-MultiLingual
MedQA-Darija-MultiLingual
The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija.
A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region.
Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.DarijaHellaSwag
Dataset Card for DarijaHellaSwag
Dataset Summary
DarijaHellaSwag is a challenging multiple-choice benchmark designed to evaluate machine reading comprehension and commonsense reasoning in Moroccan Darija. It is a translated version of the HellaSwag validation set, which presents scenarios where models must choose the most plausible continuation of a passage from four options.
Supported Tasks
Task Category: Multiple-choice question answering
Task: Answering… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaHellaSwag.Darija-SFT-Mixture
Dataset Card for Darija-SFT-Mixture
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
Darija-SFT-Mixture is a dataset consisting of 458K instruction samples, by consolidating existing Darija language resources, creating novel datasets both manually and synthetically, and translating English instructions under strict quality control.… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/Darija-SFT-Mixture.MedQA-Darija-MCQ-500
MedQA-Darija-MCQ-500 — Moroccan Darija Medical MCQ Benchmark
A 500-question multiple-choice benchmark for medical reasoning in Moroccan Arabic
Darija. To our knowledge this is the first publicly released medical MCQ benchmark in
Darija — built specifically because no equivalent existed and Darija-medical LLM
evaluation had nowhere to land.
Stat
Value
Items
500
Language
Moroccan Arabic Darija
Format
4-option MCQ (A/B/C/D, single correct)
Source stems… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedQA-Darija-MCQ-500.MoroccanHistory-QA-Darija-Dataset
Moroccan History QA Darija Dataset
This dataset is a translated subset from KBayoud/MoroccanHistory-QA-Dataset.The translation from English to Darija was done using atlasia/Terjman-Nano.The dataset is still not perfect and needs more cleaning. Please use it with caution.
Moroccan-Darija-Instruct-573K
Moroccan Darija Instruct 573K
Created by Lyte
If you use this dataset, please credit: Lyte/Moroccan-Darija-Instruct-573K
A synthetic instruction-tuning dataset of 573,175 question-answer pairs written entirely in Moroccan Darija (الدّارجة المغربية).
Dataset Summary
Metric
Value
Total rows
573,175
Original pairs
300,194
Augmented variants
272,981
Total words
~19.9M
Est. tokens
~13.3M
File size
399 MB
MSA contamination
0%
Exact duplicates
0%… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/Moroccan-Darija-Instruct-573K.Moroccan-Darija-QA
Moroccan Darija Q&A Dataset
A comprehensive question-answer dataset in Moroccan Darija (Moroccan Arabic dialect) covering various topics of daily life, culture, and practical knowledge.
📊 Dataset Overview
This dataset contains 3,470 question-answer pairs in Moroccan Darija organized across 3 configurations:
🔗 Default: 2,026 standard Q&A pairs
🌍 Translated: 1,300 translated content pairs
🧠 Reasoning: 144 reasoning-based Q&A with thinking process
Topics… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/Moroccan-Darija-QA.Health_QA_Darija
Health QA Darija — Medical QA in Moroccan Arabic (الدارجة المغربية)
Dataset Description
A curated dataset of 8,129 medical question-answer pairs in Moroccan Darija (الدارجة المغربية). Each entry contains a patient scenario, a focused medical question, and a doctor's response — all in authentic Darija. Enriched with named medical entities (symptoms, diseases, medications, tests).
🇲🇦 First large-scale medical QA dataset in Moroccan Darija — addressing the critical gap in… See the full description on the dataset page: https://huggingface.co/datasets/Kakyoin03/Health_QA_Darija.
