datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vibevoice-quran_persian-single-speakerquran-asr-mega-corpusquran-tajweed-phonetics
The complete phonetic layer of the Quran in the riwaya of Hafs 'an
'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every
phone carrying its tajweed attribution: madd class with its transmitted
length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt,
the seventeen sifat, and the rule that produced it.
Built and maintained by Quran Lab, a waqf building open technology in
the service of the Quran.
How it was built and verified
Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.quran-asr-husary
Quran ASR — Husary Muallim Dataset
Description
This dataset contains Quran recitation audio files by Sheikh Mahmoud Khalil Al-Husary at 16 kHz sampling rate, with Arabic transcriptions including diacritics.
Dataset Structure
Audio files: Stored in audio/ folder (e.g., audio/001_001.wav)
Data file: manifest.json (NeMo format)
Columns:
audio_filepath: Path to audio file
text: Arabic transcription with diacritics
duration: Audio duration in seconds
speaker:… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/quran-asr-husary.Quran_tafsir
Nasq Quranic Dataset (Arabic-English Tafsir)
لتفسير القرآن الكريم تحتوي على النص القرآني كاملاً مع التفسير الميسر (بالعربي) وتفسير المختصر (بالإنجليزي).
Columns:
id: المعرف الفريد لكل آية (من 1 إلى 6236).
surah_n: رقم السورة.
ayah_n: رقم الآية داخل السورة.
surah_name_arabic: اسم السورة باللغة العربية (تم تحديثه لضمان الدقة).
surah_name_english: اسم السورة باللغة الإنجليزية.
ayah_text_ar: نص الآية.
ayah_text_en: ترجمة نص الآية للإنجليزية.
arabic_tafsir: التفسير الميسر.… See the full description on the dataset page: https://huggingface.co/datasets/Nasaq-GP/Quran_tafsir.myanmar_quran_parallel_dataset_human_vs_ai
Myanmar Quran Parallel Dataset: Human vs AI
This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses.
It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context.
Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.itqan-quranlab-assetsQuran-Tafseers
Model Details
Developed by: Prince Sultan University - Riotu Lab
This dataset is intended for use in natural language processing tasks, particularly for understanding classical Arabic and religious texts, including text analysis, language modeling, and thematic studies.
Primary Users: Researchers and developers in the field of natural language processing, religious studies, and AI, specifically those working with classical Arabic texts.
Out-of-scope Use Cases: This dataset is not… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Quran-Tafseers.Pashto-Quran-Native-Reasoning-Dataset
Pashto-Quran-Native-Reasoning-Dataset
A specialized Pashto dataset designed for Quranic understanding, native reasoning, and natural conversational responses.
Overview
Pashto-Quran-Native-Reasoning-Dataset contains Quran-focused conversational training examples in Pashto.
The dataset is designed to help language models learn to:
understand Quranic text and its Pashto meaning
reason about the supplied content naturally
distinguish between text, translation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Quran-Native-Reasoning-Dataset.quran-multi-translator-zh
📖 Quran Chinese Multilingual NLP Corpus (High-Density RAG & Fine-Tuning Dataset)
🌟 Dataset Overview | 数据集总览
This is an elite-tier, highly-structured, and Generative Engine Optimization (GEO) focused parallel corpus for the Quran in Chinese translations. Unlike raw text scrapes, this dataset perfectly aligns the Quranic verses across 5 of the most authoritative Chinese translators, bundled with pre-calculated Knowledge Graph (KG) logic, and specifically formatted for… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/quran-multi-translator-zh.quran-burmese-word-alignment
Quran Burmese Word Alignment Dataset
Creator: freococoLicense: CC BY-NC 4.0Language: Burmese (Myanmar), ArabicFormat: JSONL (one word per line)Current Version: v10 (Surah 1–114)
📖 Overview
This dataset provides a word-by-word alignment between a Burmese (Myanmar) translation of the Quran and the original Arabic Quranic text.
Each Burmese word is represented as a single JSON object and is optionally linked to one or more corresponding Arabic word(s), with explicit… See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran-burmese-word-alignment.ARABIC_QURAN_DATASETQuran-reasoning-SFT
Quranic Reasoning Question Answering (QRQA) Dataset
Dataset Description
The Quranic Reasoning Question Answering (QRQA) Dataset is a synthetic dataset designed for experimenting purposes and for training and evaluating models capable of answering complex, knowledge-intensive questions about the Quran with a strong emphasis on reasoning. This dataset is particularly well-suited for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to enhance their understanding… See the full description on the dataset page: https://huggingface.co/datasets/musaoc/Quran-reasoning-SFT.quran-indonesia-tafseer-translationquran-tafseer-qurancom
Quran and Tafseer Dataset
Dataset Description
This dataset contains verses from the Quran along with their tafseer (interpretation/explanation).
It includes both the original Arabic text and translations, as well as detailed tafseer from scholars.
The dataset is useful for Islamic studies, NLP tasks related to religious texts, and cross-lingual research.
Dataset Structure
The dataset contains Quran verses along with their tafseer (interpretation/explanation).… See the full description on the dataset page: https://huggingface.co/datasets/gurgutan/quran-tafseer-qurancom.quranic-asr-cloud-rawdata
Quranic ASR Provider Benchmark Results
Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark.
This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio.
What Is Included
Area
Path
Purpose
Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.quran-cqa
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/sadnblueish/quran-cqa.QuranDataSetنبذة عامة عن
قاعدة بيانات القرآن الكريم
الملخص
تقدم هذه البيانات نسخة رقمية فائقة الدقة وموثقة وخالية تماماً من الأخطاء للمصحف الشريف، تم تصميمها وهندستها خصيصاً لتلبية متطلبات تدريب نماذج الذكاء الاصطناعي، معالجة اللغات الطبيعية (NLP)، وتطوير التطبيقات البرمجية واقتباس نصوص الآيات بدقة كلما كان هناك حاجة له. هذه البيانات مقدمة من "المعهد العالمي لحوسبة القرآن والعلوم الإسلامية" (ICIQIS)، ليكون بمثابة جسر يربط بين أصالة النص القرآني والقدرات التقنية الحديثة.
مصدر النص القرآني
إن النص القرآني… See the full description on the dataset page: https://huggingface.co/datasets/Khedher111/QuranDataSet.Quran_ClassificationQuran-With-Tafsirquranchapters2quran_embeddings
Quran Embeddings Dataset
This repository contains vector embeddings for the Holy Quran, generated using OpenAI's embedding model. These embeddings can be used for semantic search, question answering, and other natural language processing tasks related to Quranic text.
Dataset Information
The dataset consists of a single JSON file:
quran_embeddings.json: Contains embeddings for each verse (ayah) of the Quran with associated metadata
Metadata Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/promehedi/quran_embeddings.quranChaptersquranDataQuranVerses2wahi-quran-retrieval-pairsquran_hasanat_hadith_datasetsQuran_N_hasant_dataQuran_v2_jsonquran_hasant
