datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ArabicMMLU
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.AraDICE-ArabicMMLU-egy
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Egyptian dialect
Overview
The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic.
Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-egy.Open-ArabicaQA
ArabicaQA
ArabicaQA: Comprehensive Dataset for Arabic Question Answering
This repository contains dataset for paper ArabicaQA: Comprehensive Dataset for Arabic Question Answering. Below, we provide details regarding the materials available in this repository:
ArabicaQA is a robust dataset designed to support and advance the development of Arabic Question Answering (QA) systems. This dataset encompasses a wide range of question types, including both Machine Reading Comprehension… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/Open-ArabicaQA.documents-Egyptian-Arabic
Egyptian Arabic Mega Corpus (EAMC) — 25M Unified Egyptian Dialect Dataset
The Largest Unified Open Corpus for Egyptian Arabic (Masri / arz)
25.5M Samples | 2.66 GB (Parquet) | 9 Configs | Apache 2.0 | Ready-to-train
Comprehensive coverage: Raw Text · Wikipedia · Conversations · Speech (Whisper) · Parallel Translation (EN↔EGY) · Trilingual QA · Wikipedia Quality Classification · Fake Review / Spam Detection
Dataset Summary
Egyptian Arabic Mega Corpus… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.MedNLI
MedNLI — A Natural Language Inference Dataset For The Clinical Domain
Dataset Description
Links
Homepage:
Github.io
Repository:
Github
Paper:
arXiv
Leaderboard:
Papers with Code
Contact (Original Authors):
Alexey Romanov aromanov@cs.uml.edu, Chaitanya Shivade cshivade@us.ibm.com
Contact (Curator):
Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
`Natural Language Inference (NLI) is one of the critical tasks for… See the full description on the dataset page: https://huggingface.co/datasets/araag2/MedNLI.Shifaa_Arabic_Medical_Consultations
Shifaa Arabic Medical Consultations 🏥📊
Overview 🌍
Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses.
🔍 Why is this dataset important?
First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.arabic-qna
Sadeem QnA: An Arabic QnA Dataset 🌍✨
Welcome to the Sadeem QnA dataset, a vibrant collection designed for the advancement of Arabic natural language processing, specifically tailored for Question Answering (QnA) systems. Sourced from the rich and diverse content of Arabic Wikipedia, this dataset is a gateway to exploring the depths of Arabic language understanding, offering a unique challenge to both researchers and AI enthusiasts alike.
About Sadeem QnA
The Sadeem… See the full description on the dataset page: https://huggingface.co/datasets/sadeem-ai/arabic-qna.AraDICE-ArabicMMLU-lev
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Levantine dialect
Overview
The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic.
Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-lev.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.Arabic_EXAMS-Redux
Arabic_EXAMS-Redux
A corrected and text-repaired version of OALL/Arabic_EXAMS, the Arabic subset of the EXAMS multilingual high-school examinations benchmark.
What was fixed
Repaired corrupted Arabic text. The upstream benchmark contains widespread PDF-extraction damage to question stems and answer choices: split diacritics, fragmented words, and non-Arabic glyphs replacing standard characters. We restored these to readable Modern Standard Arabic.
Corrected the answer… See the full description on the dataset page: https://huggingface.co/datasets/inceptlabs/Arabic_EXAMS-Redux.Arabic-news-daily
Arabic News Daily 🗞️
A daily-updated, multi-domain Arabic news dataset collected automatically from 15 curated sources.
Unlike other Arabic datasets that are static snapshots, this dataset grows every day — making it ideal for research requiring fresh, current Arabic text across diverse domains.
Sources
Source
Domain
Variety
Al Jazeera Arabic
Politics
MSA
BBC Arabic
Politics
MSA
RT Arabic
Politics
MSA
Al Arabiya
Politics
MSA
AITNews
Tech & AI… See the full description on the dataset page: https://huggingface.co/datasets/unohamza/Arabic-news-daily.AraDiCE
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic.
As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.ArabCulture-Dialogue
ArabCulture-Dialogue: Cultural Benchmarking of LLMs in MSA and Arabic Dialectal Dialogue
📄 Paper (ACL 2026) | 🤗 Dataset
ArabCulture-Dialogue is the first parallel MSA–dialect cultural dialogue dataset, covering 13 Arabic-speaking countries in both Modern Standard Arabic (MSA) and each country's respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics. It contains 3,471 parallel dialogue pairs (6,942 dialogues, 343,804 words in total), each consisting… See the full description on the dataset page: https://huggingface.co/datasets/Almheiri/ArabCulture-Dialogue.Evidence_Inference_v2
Evidence Inference 2.0
Dataset Description
Links
Homepage:
Github Pages
Repository:
Github
Paper:
arXiv
Contact (Original Authors):
Jay DeYoung (deyoung.j@northeastern.edu)
Contact (Curator):
Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
The dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple treatments. Each of these articles will have multiple questions, or 'prompts'… See the full description on the dataset page: https://huggingface.co/datasets/araag2/Evidence_Inference_v2.ArabicaQA
ArabicaQA
ArabicaQA: Comprehensive Dataset for Arabic Question Answering
This repository contains dataset for paper ArabicaQA: Comprehensive Dataset for Arabic Question Answering. Below, we provide details regarding the materials available in this repository:
Dataset
Within this folder, you will find the training, validation, and test sets of the ArabicaQA dataset. Refer to the table below for the dataset statistics:
Training
Validation
Test
MRC (with answers)… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/ArabicaQA.AraDiCE-BoolQ
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the BoolQ split of the data.
Evaluation
We have used lm-harness eval framework to for the… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE-BoolQ.Shifaa_Arabic_Mental_Health_Consultations
🏥 Shifaa Arabic Mental Health Consultations 🧠
📌 Overview
Shifaa Arabic Mental Health Consultations is a high-quality dataset designed to advance Arabic medical language models.This dataset provides 35,648 real-world medical consultations, covering a wide range of mental health concerns.
📊 Dataset Summary
Size: 35,648 consultations
Main Specializations: 7
Specific Diagnoses: 123
Languages: Arabic (العربية)
Why This Dataset?
🔹 Lack of… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Mental_Health_Consultations.folkmotif
FolkMotif-270: a parallel cross-cultural entity set for cultural-bias evaluation
27 Thompson Motif-Index roles × 10 cultural traditions = 270 citation-anchored cells.
Paper: arXiv:2608.02486 · Code: github.com/AragonerUA/folkmotif
Each cell names the canonical entity that fills one structural folk-narrative role in
one tradition — the thunder-god's weapon, the smith of the gods, the goddess of love —
together with its native-script form, attested spelling variants, and a… See the full description on the dataset page: https://huggingface.co/datasets/Aragoner/folkmotif.PubMedQA
PubMedQA - A Dataset for Biomedical Research Question Answering
Dataset Description
Links
Homepage:
Github.io
Repository:
Github
Paper:
arXiv
Leaderboard:
PapersWithCode
Contact (Original Authors):
Qiao Jin (qiaojin.andy@gmail.com)
Contact (Curator):
Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial… See the full description on the dataset page: https://huggingface.co/datasets/araag2/PubMedQA.Arabic_Function_Calling
Arabic Function Calling Dataset (50K+ Samples)
مجموعة بيانات استدعاء الدوال العربية
أول وأكبر مجموعة بيانات عربية متخصصة في استدعاء الدوال (Function Calling) تغطي جميع اللهجات العربية الرئيسية والمجالات الحياتية المهمة.
Dataset Description
This is the first comprehensive Arabic function calling dataset designed for training and evaluating LLMs on Arabic tool use capabilities. The dataset covers:
5 Arabic Dialects: MSA (Modern Standard Arabic), Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_Function_Calling.AraLingBench
AraLingBench
📄 Paper: arXiv:2511.14295💻 GitHub: hammoudhasan/AraLingBench
AraLingBench is a 150-question Arabic multiple-choice benchmark that tests core linguistic competence of language models across five pillars:
النحو (Grammar)
الصرف (Morphology)
الإملاء (Spelling & Orthography)
فهم اللغة (Reading Comprehension)
التركيب اللغوي والأسلوبي (Syntax & Stylistics)
All questions are human-authored and validated, with a single correct answer and a difficulty label: Easy, Medium, or… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/AraLingBench.ArabicRAGB
ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark
Dataset Description
ArabicRAGB is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems on Arabic language tasks. Each record contains a query-passage pair where the query is grounded in the passage content.
Key Features
Passage-Grounded Queries: Each query is generated from and answerable by its paired passage
Multi-Dialect Coverage: MSA, Egyptian, Gulf… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/ArabicRAGB.saudipedia-arabic-qa
Saudipedia Q&A Dataset
Dataset Description
Summary
This dataset contains question-answer pairs scraped from Saudipedia, a comprehensive Arabic encyclopedia focused on Saudi Arabia. The dataset includes 1,082 Q&A entries covering various topics related to Saudi culture, history, economy, government, society, geography, religion, and notable personalities.
The data was collected by scraping the website's question-answer section, which provides detailed answers to… See the full description on the dataset page: https://huggingface.co/datasets/AhmadHakami/saudipedia-arabic-qa.AISA-ArabicFC
AISA-ArabicFC
Arabic Function Calling for Agentic AI Systems
The first open benchmark for tool-use in Arabic — across five dialects, eight real-world domains, and 27 structured tools.
12,125 queries · 5 dialects · 8 domains · 27 tools · 12K reasoning traces
📅 Test set releases July 20, 2026 · 🏛️ Budapest · Oct 24–29, 2026
🆕 Update — Data v1.4 & fair scoring (June 2026)
Argument scoring is now robust to surface form. A correct… See the full description on the dataset page: https://huggingface.co/datasets/TuwaiqAcademy/AISA-ArabicFC.AraTrust_undiac
AraTrust
This repository provides a modified version of the AraTrust dataset originally introduced in:
Emad A. Alghamdi, Reem I. Masoud, Deema Alnuhait, Afnan Y. Alomairi, Ahmed Ashraf, and Mohamed Zaytoon. 2025. AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic. In Proceedings of the 31st International Conference on Computational Linguistics, pages 8664–8679, Abu Dhabi, UAE. Association for Computational Linguistics.
We release a version of AraTrust dataset we used in… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/AraTrust_undiac.KazakhLawCorpus
Current Release
Current version contains three datasets.
data/
├── laws_metadata.csv
├── law_history.csv
└── law_references.csv
Dataset Description
1. laws_metadata.csv
Contains metadata describing legal acts.
Current size:
223,245 legal acts
Main fields include:
Column
Description
source_id
Internal database identifier
law_id
Stable legal act identifier
title
Original title
title_kk
Kazakh title
title_ru
Russian title… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus.SemEval_NLI4CT
NLI4CT: Multi-Evidence Natural Language Inference for Clinical Trial Reports and SemEval-2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials
Dataset Description
Links
Homepage:
sites.google
Repository:
Github2024
Paper:
arXiv2023 / arXiv2024
Leaderboard:
Codalab2023
Contact (Original Authors):Maël Jullien (mael.jullien@postgrad.manchester.ac.uk)
Contact (Curator):
Artur Guimarães (artur.guimas@gmail.com)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/araag2/SemEval_NLI4CT.HINT
HINT: Hierarchical interaction network for clinical-trial-outcome predictions
Dataset Description
Links
Homepage:
Github.io
Repository:
Github
Paper:
arXiv
Contact (Original Authors):
Tianfan Fu (futianfan@gmail.com)
Contact (Curator):Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
Clinical trials are crucial for drug development but are time consuming, expensive, and often burdensome on patients. More importantly, clinical… See the full description on the dataset page: https://huggingface.co/datasets/araag2/HINT.
