datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ArabicMMLU
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.Arabic_Aya
Dataset Card for : Arabic Aya (2A)
Arabic Aya (2A) : A Curated Subset of the Aya Collection for Arabic Language Processing
Dataset Sources & Infos
Data Origin: Derived from 69 subsets of the original Aya datasets : CohereForAI/aya_collection, CohereForAI/aya_dataset, and CohereForAI/aya_evaluation_suite.
Languages: Modern Standard Arabic (MSA) and a variety of Arabic dialects ( 'arb', 'arz', 'ary', 'ars', 'knc', 'acm', 'apc', 'aeb', 'ajp', 'acq' )… See the full description on the dataset page: https://huggingface.co/datasets/2A2I/Arabic_Aya.arabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.droid_1.0.1_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:95658"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1_test.ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.ArA-DF-2026
ArA-DF-2026
ArA-DF-2026 is an Arabic speech deepfake detection dataset for binary audio classification.
0: spoofed or synthetic speech
1: bona fide speech
The public train and dev splits include labels. The original challenge-era development-test and final-test configs remain unlabeled for reproducibility, and post-challenge gold labels are now available through the *_labeled configs.
All released audio is 16 kHz mono PCM audio packaged as lossless FLAC inside WebDataset TAR… See the full description on the dataset page: https://huggingface.co/datasets/ArabicSpeech/ArA-DF-2026.smolkalam-arabic-conversational-sft
SmolKalam
SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets.
Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.AlGhafa-Arabic-LLM-Benchmark-TranslatedAraMix-Translation-Scores
AraMix-Translation-Scores
AdaMLLab/AraMix (minhash_deduped
subset, 178,883,241 rows) with a machine-translation-detection score added to every
document. All original columns are preserved.
Columns
column
type
description
id
string
unchanged from AraMix
source
string
unchanged from AraMix
text
string
unchanged from AraMix
mmbert_quality_score
float64
AraMix's original mmbert_score, renamed
mmbert_translated_score
float64
new —… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Translation-Scores.fineweb-2-arb_ArabAraMix-Native
AraMix-Native
A native-Arabic-filtered version of
AdaMLLab/AraMix (minhash_deduped),
derived from SultanR/AraMix-Translation-Scores:
machine-translated and garbled-MT documents removed, 162,887,010 rows kept of
178,883,241 (91.06%). All columns preserved.
Filter rules
A document is kept iff all of:
mmbert_translated_score < 0.1, or a classical-text rescue: diacritic
(tashkeel) ratio ≥ 0.02 over Arabic letters and ≥ 3 distinct diacritic
classes (fully/partially… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Native.Arabic_Aya
Dataset Card for : Arabic Aya (2A)
Arabic Aya (2A) : A Curated Subset of the Aya Collection for Arabic Language Processing
Dataset Sources & Infos
Data Origin: Derived from 69 subsets of the original Aya datasets : CohereForAI/aya_collection, CohereForAI/aya_dataset, and CohereForAI/aya_evaluation_suite.
Languages: Modern Standard Arabic (MSA) and a variety of Arabic dialects ( 'arb', 'arz', 'ary', 'ars', 'knc', 'acm', 'apc', 'aeb', 'ajp'… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Arabic_Aya.FineWeb-Edu-Arabic-24M
English
العربية
FineWeb-Edu Arabic 24M
An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the sample-350BT configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics.
Highlight
Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/FineWeb-Edu-Arabic-24M.Arabic-PDtinystories_dataset_arabicArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets
ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets
Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets"
Note About Sentiment_label_confidence
"Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.arabic_chartqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_chartqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_chartqa_ar_beir.RULER-luciole_tokenizer_128k-arab-regional_v2mgap-kas_Arabdetails_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2
Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2
Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2.arabic_infographicsvqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar_beir.details_D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1arabic-univeristy-chatbot-qa
Arabic University Chatbot QA
A multilingual, multi-label intent-routing dataset for a university chatbot: given a student's
message, predict which of 20 intent categories it should route to. This is
routing, not question answering — the dataset contains no answers.
Release v0.8.0 — pinned as a Hub tag, so revision="v0.8.0" always resolves to exactly these rows.
This release holds 50,000 question rows in 22,600 scenario
groups. Every row has accepted == true; the classifier… See the full description on the dataset page: https://huggingface.co/datasets/NajahUniv/arabic-univeristy-chatbot-qa.Arabic_Aya14200
Dataset Card for : Arabic Aya (2A)
Arabic Aya (2A) : A Curated Subset of the Aya Collection for Arabic Language Processing
Dataset Sources & Infos
Data Origin: Derived from 69 subsets of the original Aya datasets : CohereForAI/aya_collection, CohereForAI/aya_dataset, and CohereForAI/aya_evaluation_suite.
Languages: Modern Standard Arabic (MSA) and a variety of Arabic dialects ( 'arb', 'arz', 'ary', 'ars', 'knc', 'acm', 'apc', 'aeb', 'ajp'… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Arabic_Aya14200.AMAZON-Products-2023-Arabic
Dataset Card for Amazon Products 2023 Arabic
Dataset Summary
This dataset contains product metadata from Amazon, filtered to include only products that became available in 2023. The dataset is intended for use in semantic search applications and includes a variety of product categories.
Number of Rows: 117,243
Number of Columns: 17
Data Source
The data is sourced from Amazon Reviews 2023.
It includes product information across multiple categories, with… See the full description on the dataset page: https://huggingface.co/datasets/milistu/AMAZON-Products-2023-Arabic.SecVulEval
Dataset Card for Dataset Name
SecVulEval is a collection of real-world C/C++ vulnerabilities.
Dataset Details
Dataset Description
The dataset is curated by collecting C/C++ vulnerability from NVD. It features statement-level vulnerable information, context information for vulnerable functions
(is_vulnerable=True), and other metadata such as CVE, CWE, commit information. The dataset contains vulnerable and non-vulnerable function samples.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/arag0rn/SecVulEval.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.droid_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 100,
"total_frames": 32212,
"total_tasks": 47,
"total_videos": 300,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:100"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_100.details_sambanovasystems__SambaLingo-Arabic-Chat-70B
Dataset Card for Evaluation run of sambanovasystems/SambaLingo-Arabic-Chat-70B
Dataset automatically created during the evaluation run of model sambanovasystems/SambaLingo-Arabic-Chat-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_sambanovasystems__SambaLingo-Arabic-Chat-70B.
