datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
canonical-islamic-corpus
🕌 Canonical Islamic Corpus (Quran + Hadith)
Description
Comprehensive corpus of authentic Islamic texts:
6,236 verses of the Holy Quran from Tanzil (Simple Clean)
315,913 unique hadith matns from four curated collectionsTotal: 322,149 texts with rich metadata.
Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection).
📊 Corpus Statistics
Metric
Value
Total entries
322,149
Quran verses
6… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/canonical-islamic-corpus.arabic-wikipedia-wikibooks-corpus
Arabic Wiki Corpus - Parquet Format
Dataset Description
This dataset contains CLEANED Arabic text from Wikipedia and Wikibooks in Parquet format (efficient, fast, columnar storage).
Key Features
✅ Parquet format - faster loading, smaller size, columnar storage
✅ Fully cleaned - no markup, no HTML, no references
✅ Arabic normalized - alef, ya, ta marbuta normalized
✅ Ready for ML/NLP/LLM - use directly without preprocessing
Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-wikipedia-wikibooks-corpus.holy-quran
🕌 The Holy Quran — Tanzil Simple Clean
Description
Complete text of the Holy Quran (6,236 ayahs) in Parquet format.Source: Tanzil Project — Simple Clean text.Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection).
📊 Dataset Statistics
Metric
Value
Total ayahs
6,236
Total surahs
114
Total juz
30
Total words
82,627
Total letters
332,837
Avg words/ayah
13.2
Avg letters/ayah
53.4… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/holy-quran.arabic-nlp-corpus
Dataset Card for Arabic NLP Corpus
Legal Notice & Rights
⚠️ Important Legal Information
This dataset contains only bibliographic metadata (titles, abstracts, author names, affiliations, DOIs, citation counts, etc.) aggregated from publicly available sources. The copyright and intellectual property rights to the original scholarly texts (including the full text of papers) belong to their respective authors, publishers, or institutions.
ArabicNLPWorld does not claim… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-nlp-corpus.
