datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.context-aware-arabic-to-english-model-with-register
Context-Aware Arabic Dialect Translation Dataset
This repository contains the dataset and code for the paper "Context-Aware Dialectal Arabic Machine Translation with Interactive Region and Register Selection" (Anonymous Submission).
Contents
context_aware_en_ar_v2.ipynb: The main Google Colab notebook used for training and evaluation.
balanced_dataset_ready.csv: The full augmented dataset (57,600 sentence pairs) produced by our RBDA pipeline.
train_dataset.csv: The… See the full description on the dataset page: https://huggingface.co/datasets/Arabic-NLP-2026/context-aware-arabic-to-english-model-with-register.ArabicNLPDatasetThe dataset, prepared in Arabic, includes 10.000 tests, 10.000 validations and 80000 train data.
The data is composed of customer comments and created from e-commerce sites.arabic-russian-scientific-translations
Arabic–Russian Scientific Translation Corpus
Description
This dataset provides parallel translations of scientific and medical texts from Arabic (original) and English (source) into Russian, generated by two state‑of‑the‑art language models:
Gemma 3:4B (Google)
LLaMA 3.1:8B (Meta)
The corpus is built from four established Arabic–English corpora (see Sources below) and is intended for machine translation, model evaluation, and linguistic research.… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-scientific-translations.canonical-islamic-corpus
🕌 Canonical Islamic Corpus (Quran + Hadith)
Description
Comprehensive corpus of authentic Islamic texts:
6,236 verses of the Holy Quran from Tanzil (Simple Clean)
315,913 unique hadith matns from four curated collectionsTotal: 322,149 texts with rich metadata.
Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection).
📊 Corpus Statistics
Metric
Value
Total entries
322,149
Quran verses
6… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/canonical-islamic-corpus.arabic-wikipedia-wikibooks-corpus
Arabic Wiki Corpus - Parquet Format
Dataset Description
This dataset contains CLEANED Arabic text from Wikipedia and Wikibooks in Parquet format (efficient, fast, columnar storage).
Key Features
✅ Parquet format - faster loading, smaller size, columnar storage
✅ Fully cleaned - no markup, no HTML, no references
✅ Arabic normalized - alef, ya, ta marbuta normalized
✅ Ready for ML/NLP/LLM - use directly without preprocessing
Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-wikipedia-wikibooks-corpus.arabic-russian-parallel-corpus
Arabic‑Russian Parallel Corpus
A parallel corpus for Arabic–Russian language pairs. Each record contains an Arabic sentence/phrase, its Russian translation, and the source of the pair.The dataset has been cleaned, deduplicated, and source names normalized to lowercase.
📊 Dataset Statistics
Overview
Metric
Value
Total entries
116,393
Unique Arabic strings
116,124
Unique Russian strings
116,152
Unique sources
6
Data… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-parallel-corpus.holy-quran
🕌 The Holy Quran — Tanzil Simple Clean
Description
Complete text of the Holy Quran (6,236 ayahs) in Parquet format.Source: Tanzil Project — Simple Clean text.Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection).
📊 Dataset Statistics
Metric
Value
Total ayahs
6,236
Total surahs
114
Total juz
30
Total words
82,627
Total letters
332,837
Avg words/ayah
13.2
Avg letters/ayah
53.4… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/holy-quran.arabic-nlp-corpus
Dataset Card for Arabic NLP Corpus
Legal Notice & Rights
⚠️ Important Legal Information
This dataset contains only bibliographic metadata (titles, abstracts, author names, affiliations, DOIs, citation counts, etc.) aggregated from publicly available sources. The copyright and intellectual property rights to the original scholarly texts (including the full text of papers) belong to their respective authors, publishers, or institutions.
ArabicNLPWorld does not claim… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-nlp-corpus.arabic-islamic-hallucination-synthetic
🧪 IslamicEval 2026 Synthetic Subtask 2 Dataset
Description
Synthetic training dataset for IslamicEval 2026 Subtask 2 (Hallucination Identification).Created from the Canonical Islamic Corpus by applying realistic Arabic/Islamic distortions.
Correct examples: original canonical texts.
Incorrect examples: texts with one or more realistic errors (word replacements, swaps, deletions, character changes, negation removal, number changes, etc.).
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-islamic-hallucination-synthetic.arabic-russian-translation-corpus
Dataset Card: Arabic-Russian Translation Corpus
Dataset Details
Dataset Description
This is a large-scale Arabic–Russian parallel corpus containing 15,467,945 sentence pairs aggregated from diverse sources: open subtitle and document corpora (OPUS), TED Talks, lexicographic dictionaries, religious texts (Quran, hadith collections, Bible), conversational phrasebooks, community-contributed examples, and news articles from political and diplomatic open… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-translation-corpus.
