datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.tinystories_dataset_arabicarabic-guardrail
Arabic Guardrail — 250,842 rows, 12 classes
Defensive dataset for training Arabic prompt-safety classifiers. Each row is an incoming user
message and which of 12 safety classes it belongs to.
بالعربية: مجموعة بيانات عربية لتدريب نماذج تصنّف الرسائل الواردة قبل وصولها للمساعد الذكي.
Arabic guardrails were a gap. Hugging Face searches for Arabic jailbreak / safety /
prompt-injection datasets return zero results, and the one Arabic guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-guardrail.arabic-prompt-routing
Arabic Prompt Routing — توجيه عربي صفري
233,720 rows. Each row is a text, a set of free-text categories, and which one it belongs
to. Categories are arbitrary Arabic — the point is a model that routes into a label set it has
never seen.
بالعربية: مجموعة بيانات عربية لتوجيه النصوص إلى فئات يكتبها المستخدم بلغة طبيعية.
الفئات ليست ثابتة، والهدف نموذج يوجّه إلى فئات لم يرها أثناء التدريب.
Arabic counterpart to the task in
LiquidAI/LFM2.5-Encoder-350M-Prompt-Router.
split… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-prompt-routing.AraMath
🗂️ Dataset: AraMath
📋 Description:
AraMath comprises 605 multiple-choice questions (MCQs) adapted from ArMath (Alghamdi et al., 2022), a dataset of mathematical word problems where each solution corresponds to a solvable equation. We reformulated the original problems into MCQ format to support structured evaluation.
🌐 Language(s): Arabic | 🧠 Task Category: Multiple Choice Questions
🔗 Hugging Face Usage Example
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/humain-ai/AraMath.arabic-corpus-audit
Arabic Corpus Integrity Audit
Author: Syamjith NK
Date: 9 September 2026 · corrected 13 September 2026
Tool: arabic-lint 0.5.0
Correction, 13 September 2026. An earlier version of this card said the labels in
Yousefmd/arabic_ocr_dataset were stored in visual order, and called that the full
reshape + bidi signature. That was wrong. Only the shaping step ran; the words are in
logical order and plain NFKC recovers them. What was measured, and stands, is that the
labels store… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arabic-corpus-audit.openvla-oft-libero-calib-dataarabic-rule-checking
Arabic Rule Checking — قواعد ونصوص عربية بأحكام محسوبة
172,488 labelled (text, rule) pairs in Arabic. Each row asks one question: does this text
satisfy this rule? The answer is مطابق or مخالف.
بالعربية: مجموعة بيانات عربية للتحقق من مطابقة النصوص لقواعد مكتوبة بلغة طبيعية. كل صف
يحتوي على نص وقاعدة وحكم محسوب آليًا، وليس رأي نموذج.
split
pairs
texts
train
159,240
48,030
validation
13,248
2,002
Built from 50,062 generated Arabic texts across 12 document types… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rule-checking.arachnid_RL
___ ___ ___ ___ ___ ___
/\ \ /\ \ /\ \ /\__\ /\ \ /\ \ _____
/::\ \ /::\ \ /::\ \ /:/ / \:\ \ \:\ \ ___ /::\ \
/:/\:\ \ /:/\:\__\ /:/\:\ \ /:/ / \:\ \ \:\ \ /\__\ /:/\:\ \
/:/ /::\ \ /:/ /:/ / /:/ /::\ \ /:/ /… See the full description on the dataset page: https://huggingface.co/datasets/webxos/arachnid_RL.arabic-itsm-dataset
Arabic ITSM Dataset
A synthetic dataset of 10,000 Arabic IT support tickets, labeled with a structured 3-level ITSM taxonomy, generated using LLMs, and validated programmatically before release.
Tickets are written in Egyptian Arabic (عامية مصرية) and cover the full range of helpdesk scenarios: access issues, network problems, hardware faults, software errors, security incidents, and service requests. Arabic technical vocabulary is mixed with English terms as they naturally… See the full description on the dataset page: https://huggingface.co/datasets/albaz2000/arabic-itsm-dataset.arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.shaer-eval-raw-gpt2-small-arabic-poetry
Shaer Evaluation Results
Models: gpt2_small_arabic_poetry
Source dataset: Shaer-AI/shaer-sft-test-generations-k5
Rows: 3481
Validation passed: True
Scored rows included: True
Dataset repo: Shaer-AI/shaer-eval-raw-gpt2-small-arabic-poetry
Files
generations.jsonl: raw generation rows
generations.csv: raw generation rows in CSV
generations_scored.jsonl: raw rows plus meter/count evaluation
validation.json: validation summary
generations_scored.csv: scored rows in… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-eval-raw-gpt2-small-arabic-poetry.arabic-rag-support-25K
Arabic RAG customer-support scenarios (27,927 rows)
Synthetic Modern Standard Arabic customer-support scenarios for training small
RAG answerers, distilled from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM.
Built as the training set for oddadmix/Nawah-50M-RAG-Support.
Each row: a customer question + the knowledge-base chunks of one fictional
company (products, prices, policies, FAQ entries) + the ideal grounded agent
answer. One generation request invents one company KB and 4 QA… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-support-25K.arabic-ocr-printed-500k
Arabic Printed OCR Lines — Synthetic, 500k
A general-purpose printed Arabic text-line recognition corpus: 500,000 train + 2,000 val
line images with labels, built to fine-tune line-recognition models (PaddleOCR PP-OCR rec
CTC/MultiHead, TrOCR, etc.). Real line-crop printed-Arabic data does not exist at this scale
on the Hub, so this corpus is rendered synthetically with diverse fonts + real Arabic text and
a documented label/decoding contract.
Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/medyas/arabic-ocr-printed-500k.English_Arabic_Translation_Pairs
English · العربية
English→Arabic Technical & Reasoning Translation Dataset
High-quality English → Modern Standard Arabic translation pairs focused on
native-English educational, scientific, and reasoning content. English source
text is drawn from real corpora (FineWeb-Edu and three NVIDIA reasoning datasets);
Arabic translations are produced by DeepSeek-v4-flash under a strict
translation-only prompt that preserves notation, numbers, formulas, code, and
citations.
Pairs… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/English_Arabic_Translation_Pairs.Ara-Egy-Medical-QAarabic-rag-chat-grpo-5K
Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations)
The reinforcement-learning half of
oddadmix/arabic-rag-chat-30K:
same generator, same validator, same schema, disjoint companies. It exists
so GRPO explores fresh knowledge bases instead of taking a second pass over
material the SFT already memorised.
conversations
turns
companies
this pool
5,259
14,018
309
Company-disjointness is exact and verified: this pool shares zero
company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.arabic-medical-qa-MERGED-MAQA-MMMLU-MIarabic_alpaca_modelarabic-english-code-switching-review-annotations
Review Annotations for Arabic-English Code-Switching Speech
This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts.
The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index.
Coverage and outcomes
The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.arabic_eou_sada_dataset
Arabic EOU SADA Dataset (Saudi Dialect)
414,053 conversational Arabic utterances annotated for End-of-Utterance (EOU) detectionStrong focus on natural Saudi dialect (خليجي / نجدي / حجازي)
Task
Binary classification:
label = 1 → End of speaker turn (EOU)
label = 0 → Speaker will continue
Columns
text: Arabic transcription
label: 0 or 1
silence_after_seconds: pause duration after this segment
split: train | validation | test (already included)… See the full description on the dataset page: https://huggingface.co/datasets/LordTenson/arabic_eou_sada_dataset.LegalRAG
Kazakh Legal Text Chunks
Dataset Summary
Kazakh Legal Text Chunks is a processed corpus of official legal texts of the Republic of Kazakhstan, prepared for retrieval-augmented generation (RAG), legal information retrieval, and grounded legal question answering in the Kazakh language.
The dataset contains structure-preserving text chunks derived from publicly available legal and normative documents. It is intended for research and development in:
legal retrieval,
legal QA… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/LegalRAG.shaer-eval-raw-gpt2-medium-arabic-poetry
Raw Shaer Continuation Generations - gpt2_medium_arabic_poetry
This dataset contains cumulative raw continuation generations for gpt2_medium_arabic_poetry.
The main table is data/test.jsonl; it includes source/prompt/reference fields plus the model output in generated_text and raw_generated_text.
Current uploaded progress target: 3481 rows out of 3481.
Scored upload: 1. When scored upload is true, reward/evaluation columns are included in the main table.
Artifacts such as… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-eval-raw-gpt2-medium-arabic-poetry.Bluemoon_Top50MB_Sorted_Fixed_ja
Bluemoon_Top50MB_Sorted_Fixed_ja
SicariusSicariiStuff/Bluemoon_Top50MB_Sorted_Fixedを、GENIAC-Team-Ozaki/karakuri-lm-8x7b-chat-v0.1-awqを用いて日本語に翻訳したロールプレイ学習用データセットです。
LLMの推論にはDeepInfraというサービスを使いました。
翻訳の詳細
3-shots promptingでの翻訳
mistralのtokenizerで出力が8000トークンを超えるまで翻訳
元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。
LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了
この結果1ターン未満となったレコード(157件)を削除… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Bluemoon_Top50MB_Sorted_Fixed_ja.Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationarabic-rag-chat-30K
Arabic multi-turn RAG customer-support conversations (31,294 conversations)
Synthetic Modern Standard Arabic customer-support conversations for training
small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite
via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4
on a local vLLM server before the run was moved off-GPU. Both teachers were given
the same prompts and the same validator. Each row is one conversation of 1-5
rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.AraDICE-HellaSwag
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the HellaSwag split of the data.
Evaluation
We have used lm-harness eval framework to for… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-HellaSwag.Amazon-2022-quarterliesARAFA
ARAFA: Arabic Fact-Checking Dataset (MSA)
ARAFA is a large publicly available Arabic fact-checking dataset in Modern Standard Arabic (MSA). It contains 181,976 claim–evidence pairs labeled as Supported, Refuted, or Not Enough Information (NEI), automatically generated and validated using large language models from Arabic Wikipedia.
The dataset is designed for training and benchmarking Arabic fact-checking models. It covers a wide range of domains, including history, politics… See the full description on the dataset page: https://huggingface.co/datasets/ChristopheKhalil/ARAFA.Arabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar
