freococo/quran_multilingual_parallel
π Qurβan Multilingual Parallel Dataset (quran_multilingual_parallel) This dataset presents a clean, structurally-aligned multilingual parallel corpus of the Qurβanic text. It is intended for linguistic, computational, and cross-lingual AI applications β not only for religious interpretation. It contains over 6,200 verse-level alignments in 54 human languages, formatted in a machine-friendly .csv structure with language-specific translation fields. π§ Datasetβ¦ See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran_multilingual_parallel.
π Qurβan Multilingual Parallel Dataset (quran_multilingual_parallel)
This dataset presents a clean, structurally-aligned multilingual parallel corpus of the Qurβanic text. It is intended for linguistic, computational, and cross-lingual AI applications β not only for religious interpretation.
It contains over 6,200 verse-level alignments in 54 human languages, formatted in a machine-friendly .csv structure with language-specific translation fields.
π§ Dataset Highlights
- π 6,236 ayahs (verses)
- π 114 surahs (chapters)
- π Translations in 53 languages
- π Arabic source text included
- π’ Fully aligned, row-per-verse CSV
- β οΈ Missing translations are transparently documented
π Files Included
π§Ύ Column Format
Each row corresponds to a verse (ayah). Columns include:
π Languages Included
This dataset includes translations in the following 54 languages:
- Arabic (original)
- Albanian
- Amharic
- Azerbaijani
- Bengali
- Bosnian
- Bulgarian
- Burmese
- Chinese
- Danish
- Dutch
- English
- Filipino
- French
- Fulah
- Persian
- German
- Gujarati
- Hausa
- Hindi
- Indonesian
- Italian
- Japanese
- Javanese
- Kazakh
- Khmer
- Korean
- Kurdish
- Kyrgyz
- Malay
- Malayalam
- Norwegian
- Pashto
- Polish
- Portuguese
- Punjabi
- Russian
- Sindhi
- Sinhalese
- Somali
- Spanish
- Swahili
- Swedish
- Tajik
- Tamil
- Tatar
- Telugu
- Thai
- Turkish
- Urdu
- Uyghur
- Uzbek
- Vietnamese
- Yoruba ---
β οΈ Missing Translation Entries
While all 6,236 verses are structurally included, 79 translations were unavailable from the source site (al-quran.cc) at the time of scraping.
These translations are missing in the following (language, surah, ayah) combinations:
<details> <summary>π Click to expand full list of missing translations (79)</summary>
- azerbaijani β al-asr:1
- azerbaijani β al-asr:2
- azerbaijani β al-asr:3
- filipino β al-haaqqa:52
- german β al-baqara:220
- german β al-isra:65
- german β an-naba:3
- german β ar-rum:5
- german β az-zamar:40
- jawa β al-furqan:37
- kyrgyz β al-mulk:16
- portuguese β aal-e-imran:50
- portuguese β aal-e-imran:51
- portuguese β an-naml:31
- portuguese β an-naziat:19
- portuguese β as-saaffat:53
- portuguese β ash-shuara:17
- portuguese β nooh:11
- portuguese β nooh:12
- portuguese β nooh:13
- portuguese β nooh:14
- portuguese β nooh:15
- portuguese β nooh:16
- portuguese β nooh:17
- portuguese β nooh:18
- portuguese β nooh:19
- portuguese β nooh:20
- sindhi β ibrahim:1
- sindhi β ibrahim:2
- sindhi β ibrahim:3
- sindhi β ibrahim:4
- sindhi β ibrahim:5
- sindhi β ibrahim:6
- sindhi β ibrahim:7
- sindhi β ibrahim:8
- sindhi β ibrahim:9
- sindhi β ibrahim:10
- sindhi β ibrahim:11
- sindhi β ibrahim:12
- sindhi β ibrahim:13
- sindhi β ibrahim:14
- sindhi β ibrahim:15
- sindhi β ibrahim:16
- sindhi β ibrahim:17
- sindhi β ibrahim:18
- sindhi β ibrahim:19
- sindhi β ibrahim:20
- sindhi β ibrahim:21
- sindhi β ibrahim:22
- sindhi β ibrahim:23
- sindhi β ibrahim:24
- sindhi β ibrahim:25
- sindhi β ibrahim:26
- sindhi β ibrahim:27
- sindhi β ibrahim:28
- sindhi β ibrahim:29
- sindhi β ibrahim:30
- sindhi β ibrahim:31
- sindhi β ibrahim:32
- sindhi β ibrahim:33
- sindhi β ibrahim:34
- sindhi β ibrahim:35
- sindhi β ibrahim:36
- sindhi β ibrahim:37
- sindhi β ibrahim:38
- sindhi β ibrahim:39
- sindhi β ibrahim:40
- sindhi β ibrahim:41
- sindhi β ibrahim:42
- sindhi β ibrahim:43
- sindhi β ibrahim:44
- sindhi β ibrahim:45
- sindhi β ibrahim:46
- sindhi β ibrahim:47
- sindhi β ibrahim:48
- sindhi β ibrahim:49
- sindhi β ibrahim:50
- sindhi β ibrahim:51
- sindhi β ibrahim:52
</details>
These empty fields are left intentionally blank in the dataset. All ayahs are present in Arabic, and structural alignment is maintained.
π Data Source and Extraction
- Source: https://www.al-quran.cc/quran-translation/
- Extraction: Python-based HTML parsing with dynamic AJAX handling
- Structural validation: 114 surahs Γ 6,236 ayahs cross-checked
- Missing translations logged during post-validation
π Linguistic and Literary Significance of the Qurβan
𧬠A Living Corpus Preserved in Speech
The Qurβan is the only major historical text that has been memorized word-for-word by millions of people across generations, regions, and languages β regardless of whether they spoke Arabic natively.
- Over 1,400 years old, the Arabic text of the Qurβan remains unchanged, recited daily, and actively memorized, verse by verse.
- It is preserved not only in manuscripts but in oral transmission, making it one of the most reliably reconstructed texts in human linguistic history.
π Translated into Languages That Werenβt Yet Standardized
Several languages in this dataset β such as todayβs form of Burmese (Myanmar), Swahili, Standard Indonesian, and even Modern English β had not yet developed or reached their current standardized written form at the time the Qurβan was first revealed over 1,400 years ago.
Yet today, these communities:
- Study and translate the Qurβan with deep linguistic care
- Retain the verse structure in all translations
- Participate in a global multilingual alignment with the same source text
This makes the Qurβan a linguistic anchor across centuries, preserved in meaning and form across dozens of linguistic systems.
βοΈ Neither Prose Nor Poetry β A Unique Register
From a literary standpoint, the Qurβanβs structure is:
- Not classical prose
- Not traditional poetry
- But a distinct rhythmic and rhetorical form known for:
- Internal rhyme and parallelism
- Recurring motifs
- High semantic density and emotional resonance
This genre continues to challenge both literary analysis and computational modeling.
π€ Relevance to AI and Linguistics
The Qurβan is an ideal corpus for AI-based multilingual NLP:
- π Fully aligned across 53 languages
- π§© Rigid source structure (verse-level)
- π§ Oral memory transmission modeling
- π Cross-lingual semantic drift analysis
- π οΈ Faithful translation alignment tasks
It provides a fixed point of semantic comparison for understanding how languages represent meaning β historically and today.
πͺͺ License
- Text: Public Domain where applicable (verify by language)
- Code/scripts: MIT License
π Citation
@misc{quran_multilingual_parallel_2025,
title = {Qurβan Multilingual Parallel Dataset},
author = {freococo},
year = {2025},
url = {https://huggingface.co/datasets/freococo/quran_multilingual_parallel}
}