datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ArabicMMLU
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.AraDICE-ArabicMMLU-egy
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Egyptian dialect
Overview
The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic.
Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-egy.Open-ArabicaQA
ArabicaQA
ArabicaQA: Comprehensive Dataset for Arabic Question Answering
This repository contains dataset for paper ArabicaQA: Comprehensive Dataset for Arabic Question Answering. Below, we provide details regarding the materials available in this repository:
ArabicaQA is a robust dataset designed to support and advance the development of Arabic Question Answering (QA) systems. This dataset encompasses a wide range of question types, including both Machine Reading Comprehension… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/Open-ArabicaQA.documents-Egyptian-Arabic
Current Hub Validation Status
Dataset Server rows: 25,399,945
Dataset Server original/Parquet size: 2,758,228,707 bytes (~2.76 GB)
Default Hub configuration currently exposes one column: text
The additional configuration names listed in this card are physical source directories and are not all recognized as separate Hub configurations. Keep the default configuration until the dataset is normalized into explicit, tested splits.
Egyptian Arabic Mega Corpus (EAMC) —… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.Shifaa_Arabic_Medical_Consultations
Shifaa Arabic Medical Consultations 🏥📊
Overview 🌍
Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses.
🔍 Why is this dataset important?
First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.arabic-qna
Sadeem QnA: An Arabic QnA Dataset 🌍✨
Welcome to the Sadeem QnA dataset, a vibrant collection designed for the advancement of Arabic natural language processing, specifically tailored for Question Answering (QnA) systems. Sourced from the rich and diverse content of Arabic Wikipedia, this dataset is a gateway to exploring the depths of Arabic language understanding, offering a unique challenge to both researchers and AI enthusiasts alike.
About Sadeem QnA
The Sadeem… See the full description on the dataset page: https://huggingface.co/datasets/sadeem-ai/arabic-qna.AraDICE-ArabicMMLU-lev
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Levantine dialect
Overview
The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic.
Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-lev.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.Arabic_EXAMS-Redux
Arabic_EXAMS-Redux
A corrected and text-repaired version of OALL/Arabic_EXAMS, the Arabic subset of the EXAMS multilingual high-school examinations benchmark.
What was fixed
Repaired corrupted Arabic text. The upstream benchmark contains widespread PDF-extraction damage to question stems and answer choices: split diacritics, fragmented words, and non-Arabic glyphs replacing standard characters. We restored these to readable Modern Standard Arabic.
Corrected the answer… See the full description on the dataset page: https://huggingface.co/datasets/inceptlabs/Arabic_EXAMS-Redux.Arabic-news-daily
Arabic News Daily 🗞️
A daily-updated, multi-domain Arabic news dataset collected automatically from 15 curated sources.
Unlike other Arabic datasets that are static snapshots, this dataset grows every day — making it ideal for research requiring fresh, current Arabic text across diverse domains.
Sources
Source
Domain
Variety
Al Jazeera Arabic
Politics
MSA
BBC Arabic
Politics
MSA
RT Arabic
Politics
MSA
Al Arabiya
Politics
MSA
AITNews
Tech & AI… See the full description on the dataset page: https://huggingface.co/datasets/unohamza/Arabic-news-daily.ArabicaQA
ArabicaQA
ArabicaQA: Comprehensive Dataset for Arabic Question Answering
This repository contains dataset for paper ArabicaQA: Comprehensive Dataset for Arabic Question Answering. Below, we provide details regarding the materials available in this repository:
Dataset
Within this folder, you will find the training, validation, and test sets of the ArabicaQA dataset. Refer to the table below for the dataset statistics:
Training
Validation
Test
MRC (with answers)… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/ArabicaQA.Shifaa_Arabic_Mental_Health_Consultations
🏥 Shifaa Arabic Mental Health Consultations 🧠
📌 Overview
Shifaa Arabic Mental Health Consultations is a high-quality dataset designed to advance Arabic medical language models.This dataset provides 35,648 real-world medical consultations, covering a wide range of mental health concerns.
📊 Dataset Summary
Size: 35,648 consultations
Main Specializations: 7
Specific Diagnoses: 123
Languages: Arabic (العربية)
Why This Dataset?
🔹 Lack of… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Mental_Health_Consultations.Arabic_Function_Calling
Arabic Function Calling Dataset (50K+ Samples)
مجموعة بيانات استدعاء الدوال العربية
أول وأكبر مجموعة بيانات عربية متخصصة في استدعاء الدوال (Function Calling) تغطي جميع اللهجات العربية الرئيسية والمجالات الحياتية المهمة.
Dataset Description
This is the first comprehensive Arabic function calling dataset designed for training and evaluating LLMs on Arabic tool use capabilities. The dataset covers:
5 Arabic Dialects: MSA (Modern Standard Arabic), Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_Function_Calling.arabic-history-and-dialects
مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦
Arabic Multi-Dialect & Civilization Instruction Dataset
المؤلف: islam-alnasherA-Dev — الحساب: https://huggingface.co/ISLAM-POالإصدار: v1.0 — التاريخ: 30 أغسطس 2026 — الترخيص: CC BY 4.0اللغة: العربية (فصحى + 4 لهجات) — الصيغة: instruction / output JSONL — الحجم: 275 عينة
📌 الملخص التنفيذي
هذه المجموعة هي مورد تعليمي متخصص لتدريب وتقييم النماذج اللغوية العربية على مسارين متوازيين:… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.ArabicRAGB
ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark
Dataset Description
ArabicRAGB is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems on Arabic language tasks. Each record contains a query-passage pair where the query is grounded in the passage content.
Key Features
Passage-Grounded Queries: Each query is generated from and answerable by its paired passage
Multi-Dialect Coverage: MSA, Egyptian, Gulf… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/ArabicRAGB.saudipedia-arabic-qa
Saudipedia Q&A Dataset
Dataset Description
Summary
This dataset contains question-answer pairs scraped from Saudipedia, a comprehensive Arabic encyclopedia focused on Saudi Arabia. The dataset includes 1,082 Q&A entries covering various topics related to Saudi culture, history, economy, government, society, geography, religion, and notable personalities.
The data was collected by scraping the website's question-answer section, which provides detailed answers to… See the full description on the dataset page: https://huggingface.co/datasets/AhmadHakami/saudipedia-arabic-qa.AISA-ArabicFC
AISA-ArabicFC
Arabic Function Calling for Agentic AI Systems
The first open benchmark for tool-use in Arabic — across five dialects, eight real-world domains, and 27 structured tools.
12,125 queries · 5 dialects · 8 domains · 27 tools · 12K reasoning traces
📅 Test set releases July 20, 2026 · 🏛️ Budapest · Oct 24–29, 2026
🆕 Update — Data v1.4 & fair scoring (June 2026)
Argument scoring is now robust to surface form. A correct… See the full description on the dataset page: https://huggingface.co/datasets/TuwaiqAcademy/AISA-ArabicFC.ArabicCulturalQA
ArabicCulturalQA
ArabicCulturalQA is the first cross-dialectal Arabic cultural QA benchmark with parallel multiple-choice (MCQ) and open-ended (OEQ) formats across Modern Standard Arabic (MSA), English, Egyptian, Levantine, Gulf, and Maghrebi. Both the MCQ and OEQ test sets have been reviewed and post-edited by native speakers of each dialect.
The dataset accompanies the LREC 2026 paper "Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants" (paper page)… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ArabicCulturalQA.mizan-iraqi-arabic-benchmark
Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set
Mizan is the first comprehensive, originally-authored evaluation benchmark
for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2
public development set: 340 originally-authored, dually-reviewed items across
two tracks (MSA baseline / Iraqi) and six axes.
📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865
🏆 Live leaderboard: https://mizan-bench.onrender.com
💻 Code… See the full description on the dataset page: https://huggingface.co/datasets/nawaralseelawi/mizan-iraqi-arabic-benchmark.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.islamic-arabic-qa
Islamic Arabic Q&A Dataset
A curated Arabic instruction-tuning dataset focused on Islamic scholarship —
covering Fiqh, Fatwa, Aqeedah, Quran Sciences, and Islamic Finance.
Built to fine-tune Arabic LLMs for Islamic Q&A tasks.
Dataset Summary
Split
Samples
Train
17,944
Validation
2,101
Test
1,042
Total
21,087
Data Sources
Source
Samples
License
SahmBenchmark/fatwa-training_standardized_new
9,953
Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/islamic-arabic-qa.islamic-qa-egyptian-arabic
Egyptian Arabic Islamic QA Dataset
Dataset Description
This dataset contains 7,465 question-answer pairs in Egyptian Arabic covering comprehensive Islamic studies topics. The dataset serves as a valuable resource for developing Arabic NLP models focused on Islamic education and religious knowledge.
Key Features
Language: Egyptian Arabic (العامية المصرية)
Domain: Islamic Studies
Size: 7,465 examples
Format: Question-Answer pairs with topic categorization… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/islamic-qa-egyptian-arabic.allaM-offsec-arabic-chat-v2
Arabic Offensive Security Chat Dataset v2
High-quality category-aware bilingual Arabic/English dataset for offensive security assistants.
What's New in v2
✅ Category-aware responses: Different response structures for web vulns, DeFi, reconnaissance tools, social engineering, etc.
✅ No generic templates: Each category has specialized analysis framework
✅ No verbatim copying: Responses analyze and transform the input, not repeat it
✅ Semantic accuracy: Tools (nmap… See the full description on the dataset page: https://huggingface.co/datasets/haiderkamal23/allaM-offsec-arabic-chat-v2.2A2I-Arabic-OpenHermes-2.5-Llama-3
Dataset Card for "2A2I-Arabic-OpenHermes-2.5-Llama-3"
Dataset Sources & Infos
Data Origin: Derived from the original Arabic OpenHermes dataset : 2A2I/Arabic-OpenHermes-2.5.
Languages: Modern Standard Arabic (MSA)
Applications: Language Modeling
License: Apache-2.0
Overview
2A2I-Arabic-OpenHermes-2.5-Llama is a Llama-3 compatible dataset carefully converted from the 2A2I's Arabic-OpenHermes-2.5 collection provided by Lyte.
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/2A2I-Arabic-OpenHermes-2.5-Llama-3.Arabic-gsm8k-v2
Dataset Summary
Arabic GSM8K is an Arabic translation of the GSM8K (Grade School Math 8K) dataset, which contains high-quality linguistically diverse grade school math word problems. The original dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning, and this Arabic version aims to extend these capabilities to Arabic language models and applications.
The dataset maintains the same characteristics as the… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-gsm8k-v2.arabic-lexicons
islamlab — The Classical Arabic Lexicons
197,731 entries from 136 lexical works — the dictionaries,
the gharīb collections and the technical glossaries — cut so that the
headword is its own column and the article is its own text.
A dictionary sits in the corpus like any other book, but nobody reads one that
way. This is the same material arranged for the thing people actually do with
it: look a word up.
work
author
entries
شمس العلوم ودواء كلام العرب من الكلوم
نشوان… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/arabic-lexicons.arabic-rag-chat-30K
Arabic multi-turn RAG customer-support conversations (31,294 conversations)
Synthetic Modern Standard Arabic customer-support conversations for training
small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite
via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4
on a local vLLM server before the run was moved off-GPU. Both teachers were given
the same prompts and the same validator. Each row is one conversation of 1-5
rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.
