datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quran-tafsir
QuranLab — Multilingual Quran Tafsir Dataset
A ready-to-use collection of Quran commentaries and annotated translations,
aligned to the canonical 6,236 ayahs.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours.
The text here reaches you through the work of QuranEnc.com, Tafsir Center for Quranic Studies, Quranic Universal Library… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-tafsir.quran
Dataset Card for the Quran
Summary
The Quran with metadata, translations, and multiple Arabic text (can use specific types for embeddings, search, classification, and display). There are 126+ columns containing 43+ languages.
TODO
Add Tafsirs
Add topics/ontology
Usage
from datasets import load_dataset
ds = load_dataset("nazimali/quran", split="train")
ds
Output:
Dataset({
features: ['surah', 'ayah', 'surah-name', 'surah-total-ayas'… See the full description on the dataset page: https://huggingface.co/datasets/nazimali/quran.hadith
Dataset Card for QuranLab — Hadith & Sunnah (Ahl al-Sunnah)
A clean Ahl-al-Sunnah hadith corpus: the canonical Sunni collections (the Six Books + the
Muwaṭṭaʾ, Musnad Aḥmad, al-Dārimī, and the famous forty-collections) in Arabic plus many
translations, with grader-attributed gradings — one config per
(collection × language). This is the audio/text family's hadith modality — companion to
quranlab/quran (Qurʾan text) and
quranlab/quran-audio (recitation).
QuranLab is a… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/hadith.quran
QuranLab — Verse-Aligned Multilingual Quran Corpus
A unified, verse-aligned multilingual Quran corpus spanning 79
languages and 185 translations. Every recension and translation is a
separate config (subset), all row-aligned on the canonical 6,236-ayah
verse_key (Hafs ʿan ʿAsim reading, 114 surahs).
The corpus also contains 111 tafsir configs: verse-grain classical
and openly licensed Arabic works, plus the native-passage and verse-expanded
views of Diyanet's Turkish Kur'an Yolu… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran.islamic-llm-training
QuranLab — Qur'an and Hadith Training Mix
Training-ready data derived from the QuranLab corpora: continued-pretraining text,
grounded instruction data, preference pairs, verifiable prompts, retrieval pairs and
a held-out evaluation set — all built on the same verse and ḥadīth keys as
quranlab/quran and
quranlab/hadith.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-llm-training.quran-qcf4
QCF4 Quran Database
A developer-friendly Quran database using QCF v4 (Quran Complex Font, version 4) glyph rendering for the Hafs recitation. This dataset provides structured, page-accurate Quranic text data paired with the fonts needed to render it — exactly as it appears in the printed Madinah Mushaf.
What is QCF4?
QCF4 is a Quranic font based on the Madinah Mushaf (1441 AH), written by the calligrapher Uthman Taha and produced by the King Fahd Complex in Madinah. It is… See the full description on the dataset page: https://huggingface.co/datasets/MohamadHajjRabee/quran-qcf4.islamic-corpus-graph
QuranLab — Qur'an & Hadith Structured Corpus and Knowledge Graph
A unified, verse- and ḥadīth-aligned structured corpus for the Qur'an and the canonical Sunnah,
assembled by volunteers under the QuranLab effort. It links Qur'anic verses, multilingual
translations, classical tafsīr, word-level morphology, and ḥadīth text with normalized authenticity
grades into one consistent graph, alongside retrieval passages, grounded question–answer pairs and a
held-out evaluation set. Every… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-corpus-graph.MuslimLife
Muslim Life Knowledge Base & RAG Dataset
Contains 90 public Simplified Chinese articles from the Salaam Alykum 穆斯林生活 / Muslim Life topic, packaged as a production-ready Hugging Face dataset with Parquet splits, Markdown article files, retrieval rows, metadata indexes, and a lightweight embedding preview layer.
[!TIP]
Human Readers / 普通读者: For normal reading, open Files and versions -> content and start with content/README.md. Example article: 2744 莱麦丹不同面貌:斋月中的人、故事与信仰现场. For… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/MuslimLife.Halal-Food-In-China
Halal Food In China RAG Corpus 🕌🍜
Dataset Overview
The Halal Food In China RAG Corpus is an expertly curated, highly structured, and multi-format dataset targeting the intersection of Chinese culinary traditions, Hui Muslim history, and Islamic dietary laws (Halal). It acts as an authoritative ground-truth database to mitigate Large Language Model (LLM) hallucinations regarding minority Islamic culture in China.
[!TIP]
Human Readers: Looking for the full… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Halal-Food-In-China.Uyghur
Uyghur Knowledge Base & RAG Dataset
Contains 74 public Simplified Chinese articles from the Salaam Alykum Uyghur topic, packaged as a production-ready Hugging Face dataset with Parquet splits, Markdown article files, retrieval rows, metadata indexes, and an embedding preview layer.
[!TIP]
Human Readers / 普通读者: Looking for normal article reading instead of raw data? Open Files and versions -> content and start with content/README.md. Example article: 3337 维吾尔民族身份是原生还是现代建构. For… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Uyghur.China-Halal-Restaurant
China Halal Restaurant Dataset (RAG Optimized) 🕌
This is a rigorously formatted Chinese Halal Restaurant corpus containing 201 authentic articles and travel guides. It is explicitly optimized for Retrieval-Augmented Generation (RAG) and pure text indexing. The data was explicitly designed to pass Hugging Face's Dataset Viewer standards natively by using optimal Parquet partitioning.
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/China-Halal-Restaurant.Chinese-Muslim-Travel
☪ Chinese-Muslim-Travel: Native Chinese Muslim Travel RAG Corpus
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
Dataset Description
Chinese-Muslim-Travel is a curated RAG corpus containing 347 native Chinese articles documenting Muslim travel, halal food, mosque architecture, and Muslim community life across 20+ countries. Every… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Chinese-Muslim-Travel.quran-tafsir
QuranLab — Multilingual Quran Tafsir Dataset
A ready-to-use collection of Quran commentaries and annotated translations,
aligned to the canonical 6,236 ayahs.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours.
The text here reaches you through the work of QuranEnc.com, Tafsir Center for Quranic Studies, Quranic Universal Library… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran-tafsir.quran
QuranLab — Verse-Aligned Multilingual Quran Corpus
A unified, verse-aligned multilingual Quran corpus spanning 79
languages and 185 translations. Every recension and translation is a
separate config (subset), all row-aligned on the canonical 6,236-ayah
verse_key (Hafs ʿan ʿAsim reading, 114 surahs).
The corpus also contains 111 tafsir configs: verse-grain classical
and openly licensed Arabic works, plus the native-passage and verse-expanded
views of Diyanet's Turkish Kur'an Yolu… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran.Hui-Muslims
Hui Muslims RAG Dataset
Contains 232 native Chinese articles.
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
TibetanMuslims
Tibetan Muslims Knowledge Base & RAG Dataset
Contains 22 public Simplified Chinese articles from the Salaam Alykum Tibetan Muslims topic, packaged as a production-ready Hugging Face dataset with Parquet splits, Markdown article files, retrieval rows, metadata indexes, and a lightweight embedding preview layer.
[!TIP]
Human Readers / 普通读者: Looking for normal article reading instead of raw data? Open Files and versions -> content and start with content/README.md. Example article:… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/TibetanMuslims.quran
Dataset Card for the Quran
Summary
The Quran with metadata, translations, and multiple Arabic text (can use specific types for embeddings, search, classification, and display). There are 126+ columns containing 43+ languages.
TODO
Add Tafsirs
Add topics/ontology
Usage
from datasets import load_dataset
ds = load_dataset("nazimali/quran", split="train")
ds
Output:
Dataset({
features: ['surah', 'ayah', 'surah-name', 'surah-total-ayas'… See the full description on the dataset page: https://huggingface.co/datasets/bakir11999/quran.islamic-articles-corpus
☪ Islamic Articles Corpus - English RAG Dataset
Dataset Description
Islamic Articles Corpus is a curated English-language RAG corpus containing 33 articles covering Muslim travel guides, mosque visits, halal food, prayer room directories, and Islamic community documentation. Every article preserves complete full-text content with all 608 embedded image references. Content focuses heavily on Singapore, Iran, Japan, Oman, and Qatar mosque and travel documentation.… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/islamic-articles-corpus.QuranExeThis dataset contains the exegeses/tafsirs (تفسير القرآن) of the holy Quran in arabic by 8 exegetes.
This is a non Official dataset. It have been scrapped from the Quran.com Api
This dataset contains 49888 records with +14 Million words. 8 records per Quranic verse
Usage Example :
from datasets import load_dataset
tafsirs = load_dataset("mustapha/QuranExe")
myanmar_quran_parallel_dataset_human_vs_ai
Myanmar Quran Parallel Dataset: Human vs AI
This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses.
It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context.
Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.quran-parallel-corpus
Quran Parallel Corpus
Verse-aligned Quran parallel corpus — Arabic (Uthmani), English (Sahih International), and Indonesian (Ministry of Religious Affairs).
Stats
Total verses: 6236
Languages: Arabic, English, Indonesian
Translation pairs: Arabic↔English, Arabic↔Indonesian, English↔Indonesian
Formats: JSONL, CSV, Parquet
Structure
Each verse record contains:
Field
Description
surah_number
Chapter (1–114)
surah_name_arabic
Arabic surah… See the full description on the dataset page: https://huggingface.co/datasets/sarjukesumo/quran-parallel-corpus.Quran_tafsir
Nasq Quranic Dataset (Arabic-English Tafsir)
لتفسير القرآن الكريم تحتوي على النص القرآني كاملاً مع التفسير الميسر (بالعربي) وتفسير المختصر (بالإنجليزي).
Columns:
id: المعرف الفريد لكل آية (من 1 إلى 6236).
surah_n: رقم السورة.
ayah_n: رقم الآية داخل السورة.
surah_name_arabic: اسم السورة باللغة العربية (تم تحديثه لضمان الدقة).
surah_name_english: اسم السورة باللغة الإنجليزية.
ayah_text_ar: نص الآية.
ayah_text_en: ترجمة نص الآية للإنجليزية.
arabic_tafsir: التفسير الميسر.… See the full description on the dataset page: https://huggingface.co/datasets/Nasaq-GP/Quran_tafsir.mofaser-quran-tafsir
Mofaser — Quran Tafsir QA Dataset
243,129 records covering all 6,236 Quranic verses interpreted through 18 classical tafsir sources (8 Arabic + 10 English), formatted as a rich QA dataset for fine-tuning and instruction-tuning LLMs.
Dataset Summary
Each record pairs a Quranic verse with a tafsir explanation in a QA format with varied question phrasings. The dataset includes full metadata: surah/ayah identifiers, Arabic and English verse text, Meccan/Medinan classification… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/mofaser-quran-tafsir.quran-multi-translator-zh
📖 Quran Chinese Multilingual NLP Corpus (High-Density RAG & Fine-Tuning Dataset)
🌟 Dataset Overview | 数据集总览
This is an elite-tier, highly-structured, and Generative Engine Optimization (GEO) focused parallel corpus for the Quran in Chinese translations. Unlike raw text scrapes, this dataset perfectly aligns the Quranic verses across 5 of the most authoritative Chinese translators, bundled with pre-calculated Knowledge Graph (KG) logic, and specifically formatted for… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/quran-multi-translator-zh.quran-burmese-word-alignment
Quran Burmese Word Alignment Dataset
Creator: freococoLicense: CC BY-NC 4.0Language: Burmese (Myanmar), ArabicFormat: JSONL (one word per line)Current Version: v10 (Surah 1–114)
📖 Overview
This dataset provides a word-by-word alignment between a Burmese (Myanmar) translation of the Quran and the original Arabic Quranic text.
Each Burmese word is represented as a single JSON object and is optionally linked to one or more corresponding Arabic word(s), with explicit… See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran-burmese-word-alignment.Chinese-Muslim
☪ Chinese-Muslim: Comprehensive Chinese Muslim Knowledge Base & RAG Corpus
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
Dataset Description
Chinese-Muslim is the flagship dataset of the Salaamalykum project — a comprehensive knowledge base containing 279 articles covering the full spectrum of Chinese Muslim life: culture, travel… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Chinese-Muslim.Pashto-Quran-Native-Reasoning-Dataset
Pashto-Quran-Native-Reasoning-Dataset
A specialized Pashto dataset designed for Quranic understanding, native reasoning, and natural conversational responses.
Overview
Pashto-Quran-Native-Reasoning-Dataset contains Quran-focused conversational training examples in Pashto.
The dataset is designed to help language models learn to:
understand Quranic text and its Pashto meaning
reason about the supplied content naturally
distinguish between text, translation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Quran-Native-Reasoning-Dataset.quran-bil-quran-connections
Dataset Card for quran-bil-quran-connections
This dataset contains a .jsonl file, where each line is a single connection between 2 verses made by Ibn Ashur in his Exegesis.
Language(s) (NLP): Arabic, English
Dataset Sources [optional]
Source Exegesis: [Tafsir al-Tahrir wa al-Tanwir]
Writeup: [Ibn Ashur's Qur’an bi’l Qur’an Visualized]
Demo: [App]
Method
For each verse/group in original source exegesis file, find all verse references using libraries like… See the full description on the dataset page: https://huggingface.co/datasets/ShahamFarooq/quran-bil-quran-connections.tibyan-quran-complete
Tibyan Quran Complete Dataset
Complete Quran dataset with 114 surahs and 6,236 ayahs in Uthmani Arabic script, with metadata and multiple text formats.
Data Fields
Field
Type
Description
surah_id
int
Surah number (1-114)
surah_name_ar
string
Arabic name of the surah
surah_name_en
string
English name of the surah
ayah_number
int
Verse number within surah
text_uthmani
string
Uthmani script (official)
text_simple
string
Simplified Arabic
juz
int
Juz… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/tibyan-quran-complete.Quran-reasoning-SFT
Quranic Reasoning Question Answering (QRQA) Dataset
Dataset Description
The Quranic Reasoning Question Answering (QRQA) Dataset is a synthetic dataset designed for experimenting purposes and for training and evaluating models capable of answering complex, knowledge-intensive questions about the Quran with a strong emphasis on reasoning. This dataset is particularly well-suited for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to enhance their understanding… See the full description on the dataset page: https://huggingface.co/datasets/musaoc/Quran-reasoning-SFT.
