datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.iraqi-arabic-sales-dialogue-dataset
Iraqi Arabic Sales Dialogue Dataset
A large synthetic dataset of Iraqi (Baghdadi-based) Arabic dialogue, centered on
retail sales, haggling, and everyday conversation.
النسخة العربية متوفرة بالكامل بالأسفل — Arabic version available in full below.
What this is
210,832 template-generated conversations, of which 171,601 (81%) are exact-unique
message sequences, spanning 20 topical categories in colloquial Iraqi Arabic. The
core of the dataset (10 categories) is… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi-arabic-sales-dialogue-dataset.AraDICE-ArabicMMLU-egy
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Egyptian dialect
Overview
The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic.
Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-egy.tinystories_dataset_arabicLlamaLens-Arabic
LlamaLens: Specialized Multilingual LLM Dataset
Overview
LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi.
LlamaLens
This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation.
Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Arabic.AraDICE-ArabicMMLU-lev
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Levantine dialect
Overview
The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic.
Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-lev.arabic-to-code-8-langs-3m
Current Hub Validation Status
Repository claim: 3,000,000 records
Dataset Server indexed rows: 1,239,045
Dataset Server estimate: 1,995,159
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive record count.
Arabic-to-Code Dataset - 8 Languages
Train an Arabic-speaking code model with a raw target of 3,000,000 Arabic instruction-to-code pairs across 8 programming… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.alpaca-gpt4-arabicThe dataset is used in the research related to MultilingualSIFT.
LlamaLens-Arabic-Native
LlamaLens: Specialized Multilingual LLM Dataset
Overview
LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi.
LlamaLens
This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation.
Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Arabic-Native.arabic-guardrail
Arabic Guardrail — 250,842 rows, 12 classes
Defensive dataset for training Arabic prompt-safety classifiers. Each row is an incoming user
message and which of 12 safety classes it belongs to.
بالعربية: مجموعة بيانات عربية لتدريب نماذج تصنّف الرسائل الواردة قبل وصولها للمساعد الذكي.
Arabic guardrails were a gap. Hugging Face searches for Arabic jailbreak / safety /
prompt-injection datasets return zero results, and the one Arabic guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-guardrail.Adala_Fonction_Publique_Maroc_Arabic_json_dataset
🏛️ Adala Fonction Publique Maroc Arabic Dataset
This dataset contains structured legal data extracted from Moroccan public service law texts, sourced from adala.justice.gov.ma. The content is in Arabic and is designed to support NLP and AI applications in legal tech, especially for Moroccan administrative and public law.
📂 Dataset Structure
Format: JSON
Language: Arabic (Standard & Legal dialect)
Content: Articles, chapters, titles from Moroccan public law texts… See the full description on the dataset page: https://huggingface.co/datasets/halimbahae/Adala_Fonction_Publique_Maroc_Arabic_json_dataset.saudi-arabic-laws-and-regulations-corpus
Saudi Arabic Laws and Regulations Corpus
A structured, article-level Arabic legal corpus containing 22,593 records from 578 official Saudi legal documents, prepared for Arabic legal information retrieval, Retrieval-Augmented Generation (RAG), grounded generation, and LLM evaluation.
Release: 1.0.0Language: ArabicDomain: Saudi laws and regulationsGranularity: Legal articleTotal records: 22,593Archival DOI: 10.5281/zenodo.21265180
Quick Start
The corpus is… See the full description on the dataset page: https://huggingface.co/datasets/SaudiArabicLaws/saudi-arabic-laws-and-regulations-corpus.Arabic_Function_Calling
Arabic Function Calling Dataset (50K+ Samples)
مجموعة بيانات استدعاء الدوال العربية
أول وأكبر مجموعة بيانات عربية متخصصة في استدعاء الدوال (Function Calling) تغطي جميع اللهجات العربية الرئيسية والمجالات الحياتية المهمة.
Dataset Description
This is the first comprehensive Arabic function calling dataset designed for training and evaluating LLMs on Arabic tool use capabilities. The dataset covers:
5 Arabic Dialects: MSA (Modern Standard Arabic), Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_Function_Calling.arabic-history-and-dialects
مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦
Arabic Multi-Dialect & Civilization Instruction Dataset
المؤلف: islam-alnasherA-Dev — الحساب: https://huggingface.co/ISLAM-POالإصدار: v1.0 — التاريخ: 30 أغسطس 2026 — الترخيص: CC BY 4.0اللغة: العربية (فصحى + 4 لهجات) — الصيغة: instruction / output JSONL — الحجم: 275 عينة
📌 الملخص التنفيذي
هذه المجموعة هي مورد تعليمي متخصص لتدريب وتقييم النماذج اللغوية العربية على مسارين متوازيين:… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.arabic-prompt-routing
Arabic Prompt Routing — توجيه عربي صفري
233,720 rows. Each row is a text, a set of free-text categories, and which one it belongs
to. Categories are arbitrary Arabic — the point is a model that routes into a label set it has
never seen.
بالعربية: مجموعة بيانات عربية لتوجيه النصوص إلى فئات يكتبها المستخدم بلغة طبيعية.
الفئات ليست ثابتة، والهدف نموذج يوجّه إلى فئات لم يرها أثناء التدريب.
Arabic counterpart to the task in
LiquidAI/LFM2.5-Encoder-350M-Prompt-Router.
split… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-prompt-routing.arabic-quotes
Arabic Quotes Dataset (arabic_Q)
The "Arabic Quotes" dataset contains a collection of Arabic quotes along with their corresponding authors and tags. The dataset is scraped from the website "arabic-quotes.com" and provides a diverse range of quotes from various authors.
Dataset Details
Version: 1.0.0
Total Quotes: 3778
Languages: Arabic
Source: arabic-quotes.com
Dataset Structure
The dataset is provided in the JSONL (JSON Lines) format, where each line… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-quotes.Egyptian-Arabic-English-Parallel-Corpus
Egyptian Arabic-English Parallel Corpus
Author: Mohamed Abdalkader · LinkedIn · GitHub
A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation.
Dataset Structure
egyptian-arabic-english-parallel-corpus/
├── SFT/
│ ├── Train/
│ │ ├── topics/ # 1,800 individual topic JSON files
│ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.ArabicRAGB
ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark
Dataset Description
ArabicRAGB is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems on Arabic language tasks. Each record contains a query-passage pair where the query is grounded in the passage content.
Key Features
Passage-Grounded Queries: Each query is generated from and answerable by its paired passage
Multi-Dialect Coverage: MSA, Egyptian, Gulf… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/ArabicRAGB.evol-instruct-arabicThe dataset is used in the research related to MultilingualSIFT.
ACVA-Arabic-Cultural-Value-Alignment
About ArabicCulture
The ArabicCulture dataset was generated by gpt3.5 and contains 8000+ True and False questions.The dataset contains questions from 58 different areas.In the answers, "True" accounted for 59.62%, and "False" accounted for 40.38%
data-all
It contains 8000+ data, and we took 5 data from each area as few-shot data.
data-select
We asked two Arabs to judge 4000 of all the data for us, and we left data that two Arabs both thought were good. Finally… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ACVA-Arabic-Cultural-Value-Alignment.arabic-corpus-audit
Arabic Corpus Integrity Audit
Author: Syamjith NK
Date: 9 September 2026 · corrected 13 September 2026
Tool: arabic-lint 0.5.0
Correction, 13 September 2026. An earlier version of this card said the labels in
Yousefmd/arabic_ocr_dataset were stored in visual order, and called that the full
reshape + bidi signature. That was wrong. Only the shaping step ran; the words are in
logical order and plain NFKC recovers them. What was measured, and stands, is that the
labels store… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arabic-corpus-audit.sharegpt-arabicArabic ShareGPT data translated by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT.
arabic-rule-checking
Arabic Rule Checking — قواعد ونصوص عربية بأحكام محسوبة
172,488 labelled (text, rule) pairs in Arabic. Each row asks one question: does this text
satisfy this rule? The answer is مطابق or مخالف.
بالعربية: مجموعة بيانات عربية للتحقق من مطابقة النصوص لقواعد مكتوبة بلغة طبيعية. كل صف
يحتوي على نص وقاعدة وحكم محسوب آليًا، وليس رأي نموذج.
split
pairs
texts
train
159,240
48,030
validation
13,248
2,002
Built from 50,062 generated Arabic texts across 12 document types… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rule-checking.Arabic-IFEvalIFEval is the first publicly available benchmark dataset specifically designed to evaluate Arabic Large Language Models (LLMs) on instruction-following capabilities in Arabic.
The dataset includes 404 high-quality, manually verified samples covering various constraints such as linguistic patterns, punctuation rules, and formatting guidelines.
Loading the Dataset
To load this dataset in Python using the 🤗 Datasets library, run the following:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/inception42/Arabic-IFEval.saudi-arabic-cs-conversations
Saudi Arabic Customer Service Conversations — Free 100 Sample
100 synthetic multi-turn conversations in authentic Saudi Arabic dialects
Built for LLM fine-tuning, chatbot training, and Arabic NLP research
Overview
This is a free 100-conversation sample from a production-quality dataset of 50,000 Saudi Arabic customer service conversations. Every conversation is fully synthetic — no real user data — and safe for commercial use.
Each conversation simulates a… See the full description on the dataset page: https://huggingface.co/datasets/dev-hussein/saudi-arabic-cs-conversations.ArabicCulturalQA
ArabicCulturalQA
ArabicCulturalQA is the first cross-dialectal Arabic cultural QA benchmark with parallel multiple-choice (MCQ) and open-ended (OEQ) formats across Modern Standard Arabic (MSA), English, Egyptian, Levantine, Gulf, and Maghrebi. Both the MCQ and OEQ test sets have been reviewed and post-edited by native speakers of each dialect.
The dataset accompanies the LREC 2026 paper "Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants" (paper page)… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ArabicCulturalQA.arabic-reasoning-dataset-logic
Arabic Logical Reasoning Tasks Dataset (Maximum 1000 Tasks)
Overview
This dataset comprises a series of logical reasoning tasks designed to evaluate and train artificial intelligence models on understanding and generating logical inferences in the Arabic language. Each task includes a unique identifier, the task type, the task text (a question and a proposed answer), and a detailed solution that outlines the thinking steps and the final answer.
Data Format
The… See the full description on the dataset page: https://huggingface.co/datasets/beetleware/arabic-reasoning-dataset-logic.alpaca_arabicarabic-itsm-dataset
Arabic ITSM Dataset
A synthetic dataset of 10,000 Arabic IT support tickets, labeled with a structured 3-level ITSM taxonomy, generated using LLMs, and validated programmatically before release.
Tickets are written in Egyptian Arabic (عامية مصرية) and cover the full range of helpdesk scenarios: access issues, network problems, hardware faults, software errors, security incidents, and service requests. Arabic technical vocabulary is mixed with English terms as they naturally… See the full description on the dataset page: https://huggingface.co/datasets/albaz2000/arabic-itsm-dataset.mizan-iraqi-arabic-benchmark
Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set
Mizan is the first comprehensive, originally-authored evaluation benchmark
for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2
public development set: 340 originally-authored, dually-reviewed items across
two tracks (MSA baseline / Iraqi) and six axes.
📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865
🏆 Live leaderboard: https://mizan-bench.onrender.com
💻 Code… See the full description on the dataset page: https://huggingface.co/datasets/nawaralseelawi/mizan-iraqi-arabic-benchmark.
