datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-books
Arabic Books
Dataset Summary
The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer.
This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.smolkalam-arabic-conversational-sft
SmolKalam
SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets.
Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.ArabicText-Large
ArabicText-Large: High-Quality Arabic Corpus for LLM Training
Dataset Summary
ArabicText-Large is a comprehensive, high-quality Arabic text corpus comprising 743,288 articles with over 244 million words, specifically curated for Large Language Model (LLM) training and fine-tuning. This dataset represents one of the largest publicly available Arabic text collections for machine learning research.
This corpus addresses the critical shortage of high-quality Arabic NLP… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/ArabicText-Large.iraqi-arabic-sales-dialogue-dataset
Iraqi Arabic Sales Dialogue Dataset
A large synthetic dataset of Iraqi (Baghdadi-based) Arabic dialogue, centered on
retail sales, haggling, and everyday conversation.
النسخة العربية متوفرة بالكامل بالأسفل — Arabic version available in full below.
What this is
210,832 template-generated conversations, of which 171,601 (81%) are exact-unique
message sequences, spanning 20 topical categories in colloquial Iraqi Arabic. The
core of the dataset (10 categories) is… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi-arabic-sales-dialogue-dataset.dclm-pro-arabic
dclm-pro-arabic
Arabic translation of DCLM-Pro (global shards 01 and 05), translated with Seed-X-PPO-7B using greedy decoding. Documents were split into ~490-token chunks at sentence boundaries, translated, and reassembled. Each row is one complete document. A companion corpus translated with the same pipeline is available at fineweb-edu-arabic.
Details
Documents: 33,245,503 (22.7% of the two source shards, uniformly sampled)
Arabic tokens: ~93B (Seed-X… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/dclm-pro-arabic.documents-Egyptian-Arabic
Dataset evaluation: See EVALUATION.md for schema/config checks, row-count status, and the data-governance plan.
Viewer note: default is a lightweight preview; select egyptian_arabic or another source config to load the full data.
Current Hub Validation Status
Dataset Server rows: 25,399,945
Dataset Server original/Parquet size: 2,758,228,707 bytes (~2.76 GB)
Default Hub configuration currently exposes one column: text
Source-specific configurations are explicitly declared in… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.FineWeb-Edu-Arabic-24M
English
العربية
FineWeb-Edu Arabic 24M
An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the sample-350BT configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics.
Highlight
Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/FineWeb-Edu-Arabic-24M.Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.fineweb-edu-arabic
fineweb-edu-arabic
Arabic translation of FineWeb-Edu (sample/350BT subset, filtered to language_score > 0.9), translated with Seed-X-PPO-7B using greedy decoding. Documents were split into ~490-token chunks, translated, and reassembled. Each row is one complete document. A companion corpus translated with the same pipeline is available at dclm-pro-arabic.
Details
Documents: 82,840,410 (27.9% of the source subset, uniformly sampled)
Arabic tokens: ~170B (Seed-X… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/fineweb-edu-arabic.Shifaa_Arabic_Medical_Consultations
Shifaa Arabic Medical Consultations 🏥📊
Overview 🌍
Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses.
🔍 Why is this dataset important?
First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.ArabicWeb24
📚 ArabicWeb24
More than 39 billion tokens of high quality Arabic web content 🌐.
What is ArabicWeb24 ?
The ArabicWeb24 dataset consists of more than 28 billion tokens of cleaned and deduplicated Arabic web data from a customized crawl.
This was processed using the large scale data processing library datatrove.
What is being released ?
We are releasing two datasets versions:
ArabicWeb24: dataset version 1 (v1) underwent extensive processing through… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/ArabicWeb24.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.arabic-to-code-8-langs-3m
Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and language-specific quality limits.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
Current Hub Validation Status
Repository claim: 3,000,000 records
Dataset Server indexed rows: 1,239,045
Dataset Server estimate: 1,995,159
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/muhammadrizo5721/101_billion_arabic_words_dataset.arabic-billion-words
Arabic Billion Words Dataset 🌕
The Abu El-Khair Arabic News Corpus (arabic-billion-words) is a comprehensive collection of Arabic text, encompassing over five million newspaper articles. The corpus is rich in linguistic diversity, containing more than a billion and a half words, with approximately three million unique words. The text is encoded in two formats: UTF-8 and Windows CP-1256, and marked up using two markup languages: SGML and XML.
Data Example
An example… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-billion-words.arabic-dialect-corpus
Arabic Dialect Corpus
A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata.
Dataset Statistics
Metric
Value
Total Records
127,180
Total Tokens
5,802,324
Average Tokens per Record
45.62
Dialect Categories
5
Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.arabic_tashkil_dataset
Arabic Tashkil (Diacritization) Dataset 📖✨
Dataset Summary
This is a massive, high-quality, Gold-Standard dataset designed explicitly for training Arabic Automatic Diacritization (Tashkil) AI models (such as ByT5, AraT5, or Custom Transformers).
The dataset contains 1,494,228 heavily vocalized pages (~2.47 GB of data) extracted from Classical Arabic and Islamic texts sourced from Thahabi.org.
To ensure the highest possible ground-truth quality, every single page… See the full description on the dataset page: https://huggingface.co/datasets/freococo/arabic_tashkil_dataset.arabic-history-and-dialects
Dataset evaluation: See EVALUATION.md for schema checks, quality limits, and the fact-check plan.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦
[!WARNING]
This is a small educational draft. The card flags three historical claims for fact-checking; verify them before using the dataset for factual QA or training.
Arabic Multi-Dialect & Civilization… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.Arabic_Function_Calling
Arabic Function Calling Dataset (50K+ Samples)
مجموعة بيانات استدعاء الدوال العربية
أول وأكبر مجموعة بيانات عربية متخصصة في استدعاء الدوال (Function Calling) تغطي جميع اللهجات العربية الرئيسية والمجالات الحياتية المهمة.
Dataset Description
This is the first comprehensive Arabic function calling dataset designed for training and evaluating LLMs on Arabic tool use capabilities. The dataset covers:
5 Arabic Dialects: MSA (Modern Standard Arabic), Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_Function_Calling.arabic_medical_dialoguearabic-math-reasoning-synth
Arabic Math Reasoning (synthetic) — مسائل رياضيات عربية مع خطوات الحل
120,462 Arabic grade-school math word problems, each with a step-by-step derivation and a
concluding sentence. Generated with gemma-3-12b-it and Qwen3.8-27B-Uncensored-NVFP4 and
arithmetically verified — every equation the reasoning states was re-evaluated, and rows whose
own arithmetic does not check out were dropped.
generator
rows
share
gemma-3-12b-it
80,480
66.8%
Qwen3.8-27B-Uncensored-NVFP4… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-math-reasoning-synth.AISA-ArabicFC
AISA-ArabicFC
Arabic Function Calling for Agentic AI Systems
The first open benchmark for tool-use in Arabic — across five dialects, eight real-world domains, and 27 structured tools.
12,125 queries · 5 dialects · 8 domains · 27 tools · 12K reasoning traces
📅 Test set releases July 20, 2026 · 🏛️ Budapest · Oct 24–29, 2026
🆕 Update — Data v1.4 & fair scoring (June 2026)
Argument scoring is now robust to surface form. A correct… See the full description on the dataset page: https://huggingface.co/datasets/TuwaiqAcademy/AISA-ArabicFC.arabic-dialect-corpus
🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi)
Dataset Description
This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA).
Languages
Primary Dialects:
Egyptian Arabic (EG) - Cairene and regional Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/arabic-dialect-corpus.Mixed-Arabic-Dataset-Main
Dataset Card for "Mixed-Arabic-Dataset"
Mixed Arabic Datasets (MAD)
The Mixed Arabic Datasets (MAD) project provides a comprehensive collection of diverse Arabic-language datasets, sourced from various repositories, platforms, and domains. These datasets cover a wide range of text types, including books, articles, Wikipedia content, stories, and more.
MAD Repo vs. MAD Main
MAD Repo
Versatility: In the MAD Repository (MAD Repo), datasets are made… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Dataset-Main.Arabic_Quotes
Arabic Quotes Dataset
Overview
The Arabic Quotes Dataset is an open-source collection of 5900+ quotes in the Arabic language, accompanied by up to three tags for each quote.
The dataset is suitable for various Natural Language Processing (NLP) tasks, such as text classification and tagging.
Data Description
Contains 5900+ quotes with up to three associated tags per quote.
All quotes and tags are in Arabic.
Use Cases
Text Classification:… See the full description on the dataset page: https://huggingface.co/datasets/AhmedBou/Arabic_Quotes.Arabic-books-and-research-dataset
Arabic reserach and books dataset (ARABD)
This dataset is an extracted cleaned text from more than 60K word files with unique arabic texts never published before.
Dataset diversity
the dataset is diverse from all kind of islamic research: [feqh, hadeeth, tafseer, tahqeeq, ... etc], from new written research to a manuscirpts.
dataset size
the dataset was more than 11GB but after cleaning (pre-processing) it becase a straight 10GB with less noisy chars.… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Arabic-books-and-research-dataset.egyptian-arabic-fake-reviews
🕵️♂️🇪🇬 FREAD: Fake Reviews Egyptian Arabic Dataset
Author: IbrahimAmin, Ismail Fakhr, M. Waleed Fakhr, Rasha Kashef License: MIT Paper: Boosting Arabic Fake Reviews Detection by Integrating Textual and Metadata Features: A Transformer-Based Model Languages: Arabic (Egyptian Dialect)
📚 Dataset Summary
FREAD is designed for detecting fake reviews in Arabic using both textual content and behavioral metadata. It contains 60,000 reviews (50K train / 10K test)… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/egyptian-arabic-fake-reviews.arabic-proverbs-collection
Arabic Proverbs Collection
Dataset Description
Arabic Proverbs Collection is a dataset containing Arabic proverbs and their explanations. The collection covers three types of Arabic proverbs from different linguistic contexts and historical periods:
Classical Arabic proverbs.
Colloquial Arabic proverbs.
Popular and contemporary Arabic proverbs generated with artificial intelligence.
Dataset Sources
Classical Arabic Proverbs
The… See the full description on the dataset page: https://huggingface.co/datasets/tahaalselwii/arabic-proverbs-collection.
