datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
islamic-sciences
islamlab — The Islamic Sciences Corpus
The Islamic sciences other than Qur'an and hadith, as their authors wrote
them: 4,022 works by scholars who died between the
0st and the 14th Hijri century, cut along their own chapter
and biographical-entry boundaries into 1,864,389 units
(3.41 billion characters of Arabic), each carrying the volume and
page it sits on so a quotation can be cited rather than merely produced.
Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.islamic-llm-training
QuranLab — Qur'an and Hadith Training Mix
Training-ready data derived from the QuranLab corpora: continued-pretraining text,
grounded instruction data, preference pairs, verifiable prompts, retrieval pairs and
a held-out evaluation set — all built on the same verse and ḥadīth keys as
quranlab/quran and
quranlab/hadith.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-llm-training.islamic-corpus-graph
QuranLab — Qur'an & Hadith Structured Corpus and Knowledge Graph
A unified, verse- and ḥadīth-aligned structured corpus for the Qur'an and the canonical Sunnah,
assembled by volunteers under the QuranLab effort. It links Qur'anic verses, multilingual
translations, classical tafsīr, word-level morphology, and ḥadīth text with normalized authenticity
grades into one consistent graph, alongside retrieval passages, grounded question–answer pairs and a
held-out evaluation set. Every… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-corpus-graph.islamic-sciences-training
islamlab — Islamic Sciences Training Sets
Training data derived from
islamlab/islamic-sciences:
text for domain adaptation, retrieval pairs with hard negatives, and citation
questions whose answers are read out of the corpus rather than written by a
model.
Nothing here is generated. Questions come from a fixed set of templates
and every answer is a field already present in the corpus. That buys a
narrow dataset in exchange for one that cannot teach a model a fact the
sources do… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences-training.islamic-arabic-qa
Islamic Arabic Q&A Dataset
A curated Arabic instruction-tuning dataset focused on Islamic scholarship —
covering Fiqh, Fatwa, Aqeedah, Quran Sciences, and Islamic Finance.
Built to fine-tune Arabic LLMs for Islamic Q&A tasks.
Dataset Summary
Split
Samples
Train
17,944
Validation
2,101
Test
1,042
Total
21,087
Data Sources
Source
Samples
License
SahmBenchmark/fatwa-training_standardized_new
9,953
Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/islamic-arabic-qa.islamic-articles-corpus
☪ Islamic Articles Corpus - English RAG Dataset
Dataset Description
Islamic Articles Corpus is a curated English-language RAG corpus containing 33 articles covering Muslim travel guides, mosque visits, halal food, prayer room directories, and Islamic community documentation. Every article preserves complete full-text content with all 608 embedded image references. Content focuses heavily on Singapore, Iran, Japan, Oman, and Qatar mosque and travel documentation.… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/islamic-articles-corpus.Training-Ai-Islamic-Dataset
🕌 Training AI Islamic Dataset
18.7M passages from classical Islamic books spanning 1,400 years of scholarship.
Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training.
📊 Dataset Structure
collections/: Categorized Islamic passages compressed in JSONL format.
metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.persian-islamicate
islamlab — Persian Islamicate Prose
237 Persian works by authors who died between the 4th and the
13th Hijri century, segmented into 73,734 units and 127,089 retrieval
passages — 125 million characters.
Persian is the second language of Islamicate learning, and this is the part of
it that survives the same gates the Arabic corpus is held to. What survives is
chiefly historiography — Bayhaqī's Tārīkh, Mīrkhwānd's Rawḍat al-ṣafā,
Sharaf al-Dīn Yazdī's Ẓafar-nāma, Abū al-Faḍl's… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/persian-islamicate.Istilah_Maliki_Dataset
[English Version | النسخة الإنجليزية]
Official Template for Jurisprudence Terminology
This dataset is based on the following contemporary scholarly source:
Book Title: Maliki School Terminology (Istilahat al-Madhab al-Maliki)
Author: Dr. Muhammad Ibrahim Ali (Former Professor of Fiqh at Umm Al-Qura University, Makkah)
Publisher: Dar al-Buhuth for Islamic Studies and Heritage Revival
Location: Dubai, UAE
Edition: 1st Edition (2000 AD / 1421 AH)
Pages: Approx. 626 pages… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/Istilah_Maliki_Dataset.islamic-biographies
islamlab — Islamic Biographical Notices
219,364 biographical notices taken out of the ṭabaqāt, tarājim and
chronicle literature and given one row each: who the notice is about, which
work it stands in, and what that work says about him.
The Muslim scholarly tradition kept biographical records for a thousand years,
mostly so that a chain of transmission could be checked. Read at scale that
record is a prosopography — who taught whom, who lived where, who was trusted
and by whom.… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-biographies.Masadir-Maliki-Dataset
📖 Dataset Summary | ملخص القاعدة
This dataset is a bibliographic Question–Answer (QA) corpus derived from the reference workSources of Maliki Jurisprudence: Usūlan wa Furūʿan by Shaykh Abū ʿĀṣim Bashīr Ḍayf (d. 1429 AH / 2008 CE).
The dataset provides structured access to:
Core Maliki fiqh sources
Authorial lineages
Methodological classifications
Printed and manuscript works across the Islamic East and West
It is intended for Islamic studies research, bibliographic analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/Masadir-Maliki-Dataset.Islamic_Finance_QnA_eval
Islamic Finance Q&A Evaluation Dataset
Validation and test splits for evaluating models on Islamic Finance Q&A.
Dataset Structure
Format: Simple prompt-answer pairs
Validation: ~203 examples (10%)
Test: ~203 examples (10%)
Language: Arabic
Domain: Islamic finance and Sharia-compliant banking
Fields
id: Unique identifier
prompt: The question prompt
question: Original question text
answer: Ground truth answer
topic: Topic category
split:… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/Islamic_Finance_QnA_eval.fiqh-maliki-talqinAl-Talqīn: A Digitized Dataset of Mālikī Jurisprudence
قاعدة بيانات كتاب «التلقين في الفقه المالكي» – رقمنة تراث فقهي
About the Book | عن الكتاب
🇬🇧 English
Al-Talqīn (التلقين) is one of the most authoritative concise manuals in Mālikī jurisprudence. It was authored by Al-Qāḍī Abū Muḥammad ʿAbd al-Wahhāb ibn ʿAlī al-Baghdādī al-Mālikī (d. 422 AH).
The book is renowned for its precision, clarity, and systematic organization, making it a foundational reference for students and scholars of the… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/fiqh-maliki-talqin.wdb-islamic-finance-benchmark
WDB Benchmark: Western Default Bias in Islamic Finance
Dataset Description
This benchmark tests whether Large Language Models exhibit Western Default Bias (WDB) - the tendency to provide Western/conventional finance answers even when the context implies Islamic finance should be used.
The Problem
When a user in Saudi Arabia or UAE asks a financial question, they likely expect Shariah-compliant advice. However, LLMs trained predominantly on Western data may… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/wdb-islamic-finance-benchmark.islamic-scholars
islamlab — Islamic Scholars Authority File
One row per scholar for the 2,420 authors whose work the corpus
carries: the name, the death year in both calendars, the school he wrote in
where the sources state one, the sciences he worked across, and the titles we
hold.
Small table, but it is the one you need first: everything else in islamlab is
keyed by work, and this is what turns a work into a person.
Spread by Hijri century — 1c: 53 · 2c: 49 · 3c: 220 · 4c: 274 · 5c: 246 · 6c:… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-scholars.canonical-islamic-corpus
🕌 Canonical Islamic Corpus (Quran + Hadith)
Description
Comprehensive corpus of authentic Islamic texts:
6,236 verses of the Holy Quran from Tanzil (Simple Clean)
315,913 unique hadith matns from four curated collectionsTotal: 322,149 texts with rich metadata.
Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection).
📊 Corpus Statistics
Metric
Value
Total entries
322,149
Quran verses
6… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/canonical-islamic-corpus.Islamic_Finance_QnA_train
Islamic Finance Q&A Training Dataset
Training split of the Islamic Finance Q&A dataset in conversational format.
Dataset Structure
Format: Conversational (human-agent pairs)
Size: ~1,624 training examples (80% of total)
Language: Arabic
Domain: Islamic finance and Sharia-compliant banking
Usage
from datasets import load_dataset
dataset = load_dataset("SahmBenchmark/Islamic_Finance_QnA_train")
train_data = dataset['train']
# Example
example =… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/Islamic_Finance_QnA_train.Islamic-Culture
☪ Islamic-Culture: Chinese Islamic Knowledge Base & RAG Corpus
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
Dataset Description
Islamic-Culture is a curated Retrieval-Augmented Generation (RAG) corpus containing 260 native Chinese articles covering Islamic culture, heritage, architecture, halal food, and Muslim community life… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Islamic-Culture.wdb-islamic-finance-benchmark2
WDB Benchmark: Western Default Bias in Islamic Finance
Dataset Description
This benchmark tests whether Large Language Models exhibit Western Default Bias (WDB) - the tendency to provide Western/conventional finance answers even when the context implies Islamic finance should be used.
How to Load
from datasets import load_dataset
ds = load_dataset("Raniahossam33/wdb-islamic-finance-benchmark")
# Access data
for sample in ds["train"]:… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/wdb-islamic-finance-benchmark2.
