datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
islamic-scholars-prosopography
📚 المستودع الشامل لتراجم وأعلام التراث الإسلامي
Unified Islamic Scholars Prosopography Repository
مستودع موحد مفتوح المصدر يضم أضخم قواعد البيانات التراجمية الأكاديمية المهيكلة لعلماء وأعلام العالم الإسلامي وحواضر الأندلس والمشرق والمغرب، والمستخرجة من أرفع المراكز البحثية العالمية: المركز الوطني الفرنسي للبحث العلمي (CNRS / IRHT)، والمجلس الأعلى للبحث العلمي الإسباني (CSIC Granada / EEA).
🏛️ محتويات المستودع (Datasets in this Repository)… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/islamic-scholars-prosopography.islamic-sciences
islamlab — The Islamic Sciences Corpus
The Islamic sciences other than Qur'an and hadith, as their authors wrote
them: 4,022 works by scholars who died between the
0st and the 14th Hijri century, cut along their own chapter
and biographical-entry boundaries into 1,864,389 units
(3.41 billion characters of Arabic), each carrying the volume and
page it sits on so a quotation can be cited rather than merely produced.
Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.islamic-corpus-graph
QuranLab — Qur'an & Hadith Structured Corpus and Knowledge Graph
A unified, verse- and ḥadīth-aligned structured corpus for the Qur'an and the canonical Sunnah,
assembled by volunteers under the QuranLab effort. It links Qur'anic verses, multilingual
translations, classical tafsīr, word-level morphology, and ḥadīth text with normalized authenticity
grades into one consistent graph, alongside retrieval passages, grounded question–answer pairs and a
held-out evaluation set. Every… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-corpus-graph.islamic-llm-training
QuranLab — Qur'an and Hadith Training Mix
Training-ready data derived from the QuranLab corpora: continued-pretraining text,
grounded instruction data, preference pairs, verifiable prompts, retrieval pairs and
a held-out evaluation set — all built on the same verse and ḥadīth keys as
quranlab/quran and
quranlab/hadith.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-llm-training.islamic-arabic-qa
Islamic Arabic Q&A Dataset
A curated Arabic instruction-tuning dataset focused on Islamic scholarship —
covering Fiqh, Fatwa, Aqeedah, Quran Sciences, and Islamic Finance.
Built to fine-tune Arabic LLMs for Islamic Q&A tasks.
Dataset Summary
Split
Samples
Train
17,944
Validation
2,101
Test
1,042
Total
21,087
Data Sources
Source
Samples
License
SahmBenchmark/fatwa-training_standardized_new
9,953
Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/islamic-arabic-qa.islamic-articles-corpus
☪ Islamic Articles Corpus - English RAG Dataset
Dataset Description
Islamic Articles Corpus is a curated English-language RAG corpus containing 33 articles covering Muslim travel guides, mosque visits, halal food, prayer room directories, and Islamic community documentation. Every article preserves complete full-text content with all 608 embedded image references. Content focuses heavily on Singapore, Iran, Japan, Oman, and Qatar mosque and travel documentation.… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/islamic-articles-corpus.islamic-sciences-training
islamlab — Islamic Sciences Training Sets
Training data derived from
islamlab/islamic-sciences:
text for domain adaptation, retrieval pairs with hard negatives, and citation
questions whose answers are read out of the corpus rather than written by a
model.
Nothing here is generated. Questions come from a fixed set of templates
and every answer is a field already present in the corpus. That buys a
narrow dataset in exchange for one that cannot teach a model a fact the
sources do… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences-training.Training-Ai-Islamic-Dataset
🕌 Training AI Islamic Dataset
18.7M passages from classical Islamic books spanning 1,400 years of scholarship.
Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training.
📊 Dataset Structure
collections/: Categorized Islamic passages compressed in JSONL format.
metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.islamic-qa-egyptian-arabic
Egyptian Arabic Islamic QA Dataset
Dataset Description
This dataset contains 7,465 question-answer pairs in Egyptian Arabic covering comprehensive Islamic studies topics. The dataset serves as a valuable resource for developing Arabic NLP models focused on Islamic education and religious knowledge.
Key Features
Language: Egyptian Arabic (العامية المصرية)
Domain: Islamic Studies
Size: 7,465 examples
Format: Question-Answer pairs with topic categorization… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/islamic-qa-egyptian-arabic.LCQA-Islamic
📚 LCQA-Islamic: A Benchmark Dataset with Larger Context for Non-Factoid QA over Islamic Texts
Dataset Summary
This dataset provides a benchmark for non-factoid question answering over Islamic texts with an emphasis on larger context retrieval.It includes expert-curated QA pairs where the answers require reasoning across multi-sentence or paragraph-level contexts from authentic Islamic sources such as the Quran, Hadith, and Tafseer.
Designed to support long-contextual… See the full description on the dataset page: https://huggingface.co/datasets/Faiz28/LCQA-Islamic.Indo-Islamic-QAThis dataset contains query–answer pairs designed for evaluating information retrieval systems, specifically for the application SEQURAN. The values in the "answer" column in file data.csv represent the IDs of relevant verses from the file knowledge-base/knowledge_base.csv.
Each ID corresponds to a specific verse entry in file knowledge-base/knowledge_base.csv, which can be referenced to retrieve the full verse content and details. It is important to note that the answers provided may vary… See the full description on the dataset page: https://huggingface.co/datasets/ramadita/Indo-Islamic-QA.Istilah_Maliki_Dataset
[English Version | النسخة الإنجليزية]
Official Template for Jurisprudence Terminology
This dataset is based on the following contemporary scholarly source:
Book Title: Maliki School Terminology (Istilahat al-Madhab al-Maliki)
Author: Dr. Muhammad Ibrahim Ali (Former Professor of Fiqh at Umm Al-Qura University, Makkah)
Publisher: Dar al-Buhuth for Islamic Studies and Heritage Revival
Location: Dubai, UAE
Edition: 1st Edition (2000 AD / 1421 AH)
Pages: Approx. 626 pages… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/Istilah_Maliki_Dataset.Masadir-Maliki-Dataset
📖 Dataset Summary | ملخص القاعدة
This dataset is a bibliographic Question–Answer (QA) corpus derived from the reference workSources of Maliki Jurisprudence: Usūlan wa Furūʿan by Shaykh Abū ʿĀṣim Bashīr Ḍayf (d. 1429 AH / 2008 CE).
The dataset provides structured access to:
Core Maliki fiqh sources
Authorial lineages
Methodological classifications
Printed and manuscript works across the Islamic East and West
It is intended for Islamic studies research, bibliographic analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/Masadir-Maliki-Dataset.islamic-biographies
islamlab — Islamic Biographical Notices
219,364 biographical notices taken out of the ṭabaqāt, tarājim and
chronicle literature and given one row each: who the notice is about, which
work it stands in, and what that work says about him.
The Muslim scholarly tradition kept biographical records for a thousand years,
mostly so that a chain of transmission could be checked. Read at scale that
record is a prosopography — who taught whom, who lived where, who was trusted
and by whom.… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-biographies.fiqh-maliki-talqinAl-Talqīn: A Digitized Dataset of Mālikī Jurisprudence
قاعدة بيانات كتاب «التلقين في الفقه المالكي» – رقمنة تراث فقهي
About the Book | عن الكتاب
🇬🇧 English
Al-Talqīn (التلقين) is one of the most authoritative concise manuals in Mālikī jurisprudence. It was authored by Al-Qāḍī Abū Muḥammad ʿAbd al-Wahhāb ibn ʿAlī al-Baghdādī al-Mālikī (d. 422 AH).
The book is renowned for its precision, clarity, and systematic organization, making it a foundational reference for students and scholars of the… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/fiqh-maliki-talqin.maliki-terminology
اصطلاحات أعلام المالكية وألقابهم
قاعدة بيانات متخصصة في رموز وألقاب المدرسة المالكية
📖 وصف البيانات
تحتوي هذه القاعدة على استخراج دقيق للألقاب والرموز العلمية المستخدمة في كتب الفقه المالكي (مثل: الشيخ، الأخوان، الصادقان، المحمدون...). تم تحويل المادة العلمية إلى صيغة سؤال وجواب (QA) لتسهيل تدريب نماذج الذكاء الاصطناعي على فهم السياق التاريخي والعلمي للمذهب.
📂 محتوى الملف
عدد القيود: 29 اصطلاحاً رئيسياً.
الصيغة: JSONL.
الحقول: - question: السؤال… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/maliki-terminology.Islamic_Finance_QnA_eval
Islamic Finance Q&A Evaluation Dataset
Validation and test splits for evaluating models on Islamic Finance Q&A.
Dataset Structure
Format: Simple prompt-answer pairs
Validation: ~203 examples (10%)
Test: ~203 examples (10%)
Language: Arabic
Domain: Islamic finance and Sharia-compliant banking
Fields
id: Unique identifier
prompt: The question prompt
question: Original question text
answer: Ground truth answer
topic: Topic category
split:… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/Islamic_Finance_QnA_eval.islamic-historical-corpus
Islamic Historical Corpus — Source Manifest
The public source manifest for the
Islamic Historical Corpus — the world's first
classified, AI-ready Islamic historical knowledge base.
What This Dataset Contains
This dataset is the public source manifest only —
a catalog of the 216 primary sources ingested into the
corpus, with authentication tiers and chunk counts.
No copyrighted translation text is included.
File
Records
Description
sources_manifest.json
216
Full… See the full description on the dataset page: https://huggingface.co/datasets/IslamStories/islamic-historical-corpus.wdb-islamic-finance-benchmark
WDB Benchmark: Western Default Bias in Islamic Finance
Dataset Description
This benchmark tests whether Large Language Models exhibit Western Default Bias (WDB) - the tendency to provide Western/conventional finance answers even when the context implies Islamic finance should be used.
The Problem
When a user in Saudi Arabia or UAE asks a financial question, they likely expect Shariah-compliant advice. However, LLMs trained predominantly on Western data may… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/wdb-islamic-finance-benchmark.Islamic-Culture
☪ Islamic-Culture: Chinese Islamic Knowledge Base & RAG Corpus
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
Dataset Description
Islamic-Culture is a curated Retrieval-Augmented Generation (RAG) corpus containing 260 native Chinese articles covering Islamic culture, heritage, architecture, halal food, and Muslim community life… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Islamic-Culture.islamic-scholars
islamlab — Islamic Scholars Authority File
One row per scholar for the 2,420 authors whose work the corpus
carries: the name, the death year in both calendars, the school he wrote in
where the sources state one, the sciences he worked across, and the titles we
hold.
Small table, but it is the one you need first: everything else in islamlab is
keyed by work, and this is what turns a work into a person.
Spread by Hijri century — 1c: 53 · 2c: 49 · 3c: 220 · 4c: 274 · 5c: 246 · 6c:… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-scholars.wdb-islamic-finance-benchmark2
WDB Benchmark: Western Default Bias in Islamic Finance
Dataset Description
This benchmark tests whether Large Language Models exhibit Western Default Bias (WDB) - the tendency to provide Western/conventional finance answers even when the context implies Islamic finance should be used.
How to Load
from datasets import load_dataset
ds = load_dataset("Raniahossam33/wdb-islamic-finance-benchmark")
# Access data
for sample in ds["train"]:… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/wdb-islamic-finance-benchmark2.palmx_2025_subtask2_islamic
🏷️ PalmX 2025 — Islamic Culture Evaluation (PalmX-IC)
Dataset Summary
PalmX-IC assesses a model’s knowledge of Islamic culture—rituals, Qurʾān verses, Ḥadīth, historic events, jurisprudence, and religious holidays—core elements of life across the Arab world.All items are authored in Modern Standard Arabic (MSA) . The dataset powers Subtask 2 of the PalmX 2025 shared task.
Dataset Structure
Split
# MCQs
Release Date
Notes
Train
600
10 Jun 2025
With… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/palmx_2025_subtask2_islamic.uzbek-islamic-qa-v1
Uzbek Islamic Q&A (v1)
⚠️ Sensitivity & disclaimer — read first
This is an archival dataset of already-published question-and-answer content,
collected from the public website savollar.islom.uz ("Zikr ahlidan so'rang")
and restructured for natural-language-processing research.
The answers reflect the views of specific scholars — primarily the late
Shayx Muhammad Sodiq Muhammad Yusuf and the scholars associated with islom.uz —
as published on that site. They… See the full description on the dataset page: https://huggingface.co/datasets/sukhrobnurali/uzbek-islamic-qa-v1.Quranlab-islamic-dataset
🌙 QuranLab Islamic Dataset
Unified dataset for Quranic recitation analysis, Hadith verification, Fiqh QA, and Abjad validationPart of the ADANiD Ecosystem
📊 Dataset Splits
Split
Quran Recitations
Hadith Texts
Fiqh Questions
Abjad Data
Train
40K samples
20K entries
12K pairs
6,236 verses
Test
5K samples
2.5K entries
1.5K pairs
6,236 verses
Validation
5K samples
2.5K entries
1.5K pairs
6,236 verses
🧠 Use Cases
Abjad Validation:… See the full description on the dataset page: https://huggingface.co/datasets/ADANiD/Quranlab-islamic-dataset.
