datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
QuranTTS
QuranTTS v4
An ear-verified Quranic recitation corpus for TTS and speech restoration.
Built in-house for the lab's own training runs. We have since moved on to a larger corpus and a newer
pipeline, so this one is published rather than shelved. What you get is the corpus exactly as it stood when
we stopped using it: complete, documented, and unmaintained.
Non-commercial, strictly. Neither this corpus nor any model trained on it may be used for any
commercial purpose, and… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/QuranTTS.quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/hetchyy/quranic-universal-ayahs.quran-asr-mega-corpusQuranTTS
QuranTTS v4
An ear-verified Quranic recitation corpus for TTS and speech restoration.
v4 supersedes the earlier v1_chunks / v2_clean / v2_raw configurations.
It is a full re-cut of the corpus at 24-bit / 48 kHz, with ayah boundaries
taken from forced alignment rather than fixed padding — which fixes the
truncated ghunnah endings present in earlier releases.
Clips
64,721
Duration
301.4 h
Reciters
16
Riwaya / style
Hafs, murattal
Coverage
112 surahs, 6… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/QuranTTS.quran-tajweed-phonetics
The complete phonetic layer of the Quran in the riwaya of Hafs 'an
'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every
phone carrying its tajweed attribution: madd class with its transmitted
length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt,
the seventeen sifat, and the rule that produced it.
Built and maintained by Quran Lab, a waqf building open technology in
the service of the Quran.
How it was built and verified
Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.quran-tafsir
QuranLab — Multilingual Quran Tafsir Dataset
A ready-to-use collection of Quran commentaries and annotated translations,
aligned to the canonical 6,236 ayahs.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours.
The text here reaches you through the work of QuranEnc.com, Tafsir Center for Quranic Studies, Quranic Universal Library… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-tafsir.quran
Dataset Card for the Quran
Summary
The Quran with metadata, translations, and multiple Arabic text (can use specific types for embeddings, search, classification, and display). There are 126+ columns containing 43+ languages.
TODO
Add Tafsirs
Add topics/ontology
Usage
from datasets import load_dataset
ds = load_dataset("nazimali/quran", split="train")
ds
Output:
Dataset({
features: ['surah', 'ayah', 'surah-name', 'surah-total-ayas'… See the full description on the dataset page: https://huggingface.co/datasets/nazimali/quran.hadith
Dataset Card for QuranLab — Hadith & Sunnah (Ahl al-Sunnah)
A clean Ahl-al-Sunnah hadith corpus: the canonical Sunni collections (the Six Books + the
Muwaṭṭaʾ, Musnad Aḥmad, al-Dārimī, and the famous forty-collections) in Arabic plus many
translations, with grader-attributed gradings — one config per
(collection × language). This is the audio/text family's hadith modality — companion to
quranlab/quran (Qurʾan text) and
quranlab/quran-audio (recitation).
QuranLab is a… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/hadith.quran-audio-text
QuranLab — Verse-Aligned Quran Text + Recitation References
This dataset joins QuranLab's canonical Hafs Arabic text to its
per-ayah recitation references. Every row is one exact
(recitation_id, verse_key) pair: the Uthmani transcript, a search-friendly
Simple-Clean transcript, and the corresponding audio_url.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio-text.quran-audio
Dataset Card for QuranLab — Qur'an Recitation Audio
A verse-aligned reference layer for Qur'an recitation audio: a unified taxonomy of 245 reciters, per-ayah and per-surah reference manifests that link every verse to its original public source, and word-precise timing released under CC-BY-4.0. This dataset hosts no audio files — it points to each recitation where it already lives and contributes the connective scholarship: taxonomy, alignment, and timing.
QuranLab is a… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio.quran-ijaz
QuranLab — Iʿjāz, Letter Structure and Abjad Reckoning
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours.
The text these numbers are counted from reaches you through the work of the Tanzil Project, and the classical works through the OpenITI release. Who to thank, and what we consulted without ever quoting, is set… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-ijaz.quran
QuranLab — Verse-Aligned Multilingual Quran Corpus
A unified, verse-aligned multilingual Quran corpus spanning 79
languages and 185 translations. Every recension and translation is a
separate config (subset), all row-aligned on the canonical 6,236-ayah
verse_key (Hafs ʿan ʿAsim reading, 114 surahs).
The corpus also contains 111 tafsir configs: verse-grain classical
and openly licensed Arabic works, plus the native-passage and verse-expanded
views of Diyanet's Turkish Kur'an Yolu… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran.quran-qcf4
QCF4 Quran Database
A developer-friendly Quran database using QCF v4 (Quran Complex Font, version 4) glyph rendering for the Hafs recitation. This dataset provides structured, page-accurate Quranic text data paired with the fonts needed to render it — exactly as it appears in the printed Madinah Mushaf.
What is QCF4?
QCF4 is a Quranic font based on the Madinah Mushaf (1441 AH), written by the calligrapher Uthman Taha and produced by the King Fahd Complex in Madinah. It is… See the full description on the dataset page: https://huggingface.co/datasets/MohamadHajjRabee/quran-qcf4.islamic-corpus-graph
QuranLab — Qur'an & Hadith Structured Corpus and Knowledge Graph
A unified, verse- and ḥadīth-aligned structured corpus for the Qur'an and the canonical Sunnah,
assembled by volunteers under the QuranLab effort. It links Qur'anic verses, multilingual
translations, classical tafsīr, word-level morphology, and ḥadīth text with normalized authenticity
grades into one consistent graph, alongside retrieval passages, grounded question–answer pairs and a
held-out evaluation set. Every… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-corpus-graph.Quran-kabyle-ayt-mensour
Dataset Card: Quran Kabyle Translation (Ramdane At Mensour)
Dataset Summary
This dataset contains the Kabyle (Taqbaylit / Amazigh) translation of the Holy Quran titled "LEQWṚAN S TMAZIƔT", translated by Ramdane At Mensour (Remḍan At Menṣuṛ). It provides verse-by-verse alignments across three script representations: legacy custom-encoded ASCII, standardized INALCO Latin, and IRCAM Tifinagh.
Previously, digital distributions of this translation across mobile apps… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Quran-kabyle-ayt-mensour.MuslimLife
Muslim Life Knowledge Base & RAG Dataset
Contains 90 public Simplified Chinese articles from the Salaam Alykum 穆斯林生活 / Muslim Life topic, packaged as a production-ready Hugging Face dataset with Parquet splits, Markdown article files, retrieval rows, metadata indexes, and a lightweight embedding preview layer.
[!TIP]
Human Readers / 普通读者: For normal reading, open Files and versions -> content and start with content/README.md. Example article: 2744 莱麦丹不同面貌:斋月中的人、故事与信仰现场. For… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/MuslimLife.Uyghur
Uyghur Knowledge Base & RAG Dataset
Contains 74 public Simplified Chinese articles from the Salaam Alykum Uyghur topic, packaged as a production-ready Hugging Face dataset with Parquet splits, Markdown article files, retrieval rows, metadata indexes, and an embedding preview layer.
[!TIP]
Human Readers / 普通读者: Looking for normal article reading instead of raw data? Open Files and versions -> content and start with content/README.md. Example article: 3337 维吾尔民族身份是原生还是现代建构. For… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Uyghur.China-Halal-Restaurant
China Halal Restaurant Dataset (RAG Optimized) 🕌
This is a rigorously formatted Chinese Halal Restaurant corpus containing 201 authentic articles and travel guides. It is explicitly optimized for Retrieval-Augmented Generation (RAG) and pure text indexing. The data was explicitly designed to pass Hugging Face's Dataset Viewer standards natively by using optimal Parquet partitioning.
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/China-Halal-Restaurant.Chinese-Muslim-Travel
☪ Chinese-Muslim-Travel: Native Chinese Muslim Travel RAG Corpus
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
Dataset Description
Chinese-Muslim-Travel is a curated RAG corpus containing 347 native Chinese articles documenting Muslim travel, halal food, mosque architecture, and Muslim community life across 20+ countries. Every… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Chinese-Muslim-Travel.quran-tafsir
QuranLab — Multilingual Quran Tafsir Dataset
A ready-to-use collection of Quran commentaries and annotated translations,
aligned to the canonical 6,236 ayahs.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours.
The text here reaches you through the work of QuranEnc.com, Tafsir Center for Quranic Studies, Quranic Universal Library… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran-tafsir.quran-question-answer-context
Dataset Card for "quran-question-answer-context"
Dataset Summary
Translated the original dataset from Arabic to English and added the Surah ayahs to the context column.
Usage
from datasets import load_dataset
dataset = load_dataset("nazimali/quran-question-answer-context")
DatasetDict({
train: Dataset({
features: ['q_id', 'question', 'answer', 'q_word', 'q_topic', 'fine_class', 'class', 'ontology_concept', 'ontology_concept2', 'source', 'q_src_id'… See the full description on the dataset page: https://huggingface.co/datasets/nazimali/quran-question-answer-context.quran
QuranLab — Verse-Aligned Multilingual Quran Corpus
A unified, verse-aligned multilingual Quran corpus spanning 79
languages and 185 translations. Every recension and translation is a
separate config (subset), all row-aligned on the canonical 6,236-ayah
verse_key (Hafs ʿan ʿAsim reading, 114 surahs).
The corpus also contains 111 tafsir configs: verse-grain classical
and openly licensed Arabic works, plus the native-passage and verse-expanded
views of Diyanet's Turkish Kur'an Yolu… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran.Hui-Muslims
Hui Muslims RAG Dataset
Contains 232 native Chinese articles.
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
TibetanMuslims
Tibetan Muslims Knowledge Base & RAG Dataset
Contains 22 public Simplified Chinese articles from the Salaam Alykum Tibetan Muslims topic, packaged as a production-ready Hugging Face dataset with Parquet splits, Markdown article files, retrieval rows, metadata indexes, and a lightweight embedding preview layer.
[!TIP]
Human Readers / 普通读者: Looking for normal article reading instead of raw data? Open Files and versions -> content and start with content/README.md. Example article:… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/TibetanMuslims.QuranQuran Dataset with Various Translations
License: Creative Commons (CC)
Task Categories: Translation
Languages:
Arabic (ar)
English (en)
Urdu (ur)
Turkish (tr)
Spanish (es)
Swedish (sv)
Bengali (bn)
Indonesian (id)
French (fr)
Russian (ru)
Chinese (zh)
Tags: Islam, Quran, Translation
Description:
This dataset contains complete translations of the Quran in various languages. The Quran is the holy book of Islam, and this dataset provides access to its text in multiple languages, making it a… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Quran.quran
Dataset Card for the Quran
Summary
The Quran with metadata, translations, and multiple Arabic text (can use specific types for embeddings, search, classification, and display). There are 126+ columns containing 43+ languages.
TODO
Add Tafsirs
Add topics/ontology
Usage
from datasets import load_dataset
ds = load_dataset("nazimali/quran", split="train")
ds
Output:
Dataset({
features: ['surah', 'ayah', 'surah-name', 'surah-total-ayas'… See the full description on the dataset page: https://huggingface.co/datasets/bakir11999/quran.islamic-articles-corpus
☪ Islamic Articles Corpus - English RAG Dataset
Dataset Description
Islamic Articles Corpus is a curated English-language RAG corpus containing 33 articles covering Muslim travel guides, mosque visits, halal food, prayer room directories, and Islamic community documentation. Every article preserves complete full-text content with all 608 embedded image references. Content focuses heavily on Singapore, Iran, Japan, Oman, and Qatar mosque and travel documentation.… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/islamic-articles-corpus.quran-dargwa-parallel
Quran Arabic–Dargwa Parallel Corpus
A verse-aligned parallel corpus of the Quran in Arabic and Dargwa.
The dataset contains 6,236 aligned records covering all 114 surahs. Each record contains an Arabic verse and its Dargwa translation.
The Dargwa text is based on the translation by Magomed Gamidov, published by Yupiter in Makhachkala in 1995. The printed edition was digitized using OCR, corrected semi-automatically, and partially reviewed manually. A small number of OCR… See the full description on the dataset page: https://huggingface.co/datasets/Murtazali/quran-dargwa-parallel.Quran-Persian
Quran Persian Translation
Quran Persian Translation is a Persian dataset of over 20 hours of audio and text pairs designed for speech synthesis and other speech-related tasks. The dataset has been collected, processed, and annotated as a part of the Mana-TTS project. For details on data processing pipeline and statistics, please refer to the paper in the Citation secition.
Acknowledgement
The raw audio and text files have been collected from the Persian translation… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/Quran-Persian.
