ArabicNLP
ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.context-aware-arabic-to-english-model-with-register
Context-Aware Arabic Dialect Translation Dataset
This repository contains the dataset and code for the paper "Context-Aware Dialectal Arabic Machine Translation with Interactive Region and Register Selection" (Anonymous Submission).
Contents
context_aware_en_ar_v2.ipynb: The main Google Colab notebook used for training and evaluation.
balanced_dataset_ready.csv: The full augmented dataset (57,600 sentence pairs) produced by our RBDA pipeline.
train_dataset.csv: The… See the full description on the dataset page: https://huggingface.co/datasets/Arabic-NLP-2026/context-aware-arabic-to-english-model-with-register.ArabicNLPDatasetThe dataset, prepared in Arabic, includes 10.000 tests, 10.000 validations and 80000 train data.
The data is composed of customer comments and created from e-commerce sites.arabic-russian-scientific-translations
Arabic–Russian Scientific Translation Corpus
Description
This dataset provides parallel translations of scientific and medical texts from Arabic (original) and English (source) into Russian, generated by two state‑of‑the‑art language models:
Gemma 3:4B (Google)
LLaMA 3.1:8B (Meta)
The corpus is built from four established Arabic–English corpora (see Sources below) and is intended for machine translation, model evaluation, and linguistic research.… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-scientific-translations.canonical-islamic-corpus
🕌 Canonical Islamic Corpus (Quran + Hadith)
Description
Comprehensive corpus of authentic Islamic texts:
6,236 verses of the Holy Quran from Tanzil (Simple Clean)
315,913 unique hadith matns from four curated collectionsTotal: 322,149 texts with rich metadata.
Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection).
📊 Corpus Statistics
Metric
Value
Total entries
322,149
Quran verses
6… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/canonical-islamic-corpus.arabic-wikipedia-wikibooks-corpus
Arabic Wiki Corpus - Parquet Format
Dataset Description
This dataset contains CLEANED Arabic text from Wikipedia and Wikibooks in Parquet format (efficient, fast, columnar storage).
Key Features
✅ Parquet format - faster loading, smaller size, columnar storage
✅ Fully cleaned - no markup, no HTML, no references
✅ Arabic normalized - alef, ya, ta marbuta normalized
✅ Ready for ML/NLP/LLM - use directly without preprocessing
Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-wikipedia-wikibooks-corpus.
