datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zo-bible
zo-bible
A sentence-aligned parallel Bible corpus covering 8 closely related Zo speech varieties and English across 10 translation versions (30,974 canonical verses).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr), Mizo (lus), Paite (pck), Vaiphei (vap), Thadou (tcz), Gangte (gnb), Zou (zom), English (eng)
Family: Zo Languages
Volume: 30,974 verse anchors across 66 canonical books (10 translation editions)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/zo-bible.bible
The Bible in 1,004 Languages
14,497,397 verses across 1,253 translations in 1,004 languages, every verse
keyed to the same chapter-and-verse address so that any two languages can be
aligned by joining on book, chapter and verse.
The Bible is the most widely translated text in existence, and for several
hundred of the languages here it is the largest — sometimes the only —
substantial digitised text. That makes this corpus unusually useful for
low-resource machine translation… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible.bible-caucasus
Bible Translations in Languages of the Caucasus
Verse-aligned Bible translations in 22 languages of the Caucasus (plus Russian and
English as pivots). Most are published by the
Institute for Bible Translation (IBT, Moscow); the Udi edition
comes from Translation Services International; the two
English editions (WEB, KJVA) are public-domain reference translations from the same IBT
server (see Provenance). One config per language/edition, 400,286 verse rows total.
Every row… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/bible-caucasus.bible-parallel-english
Parallel Bible — English Translations and Ancient Versions
A verse-aligned parallel corpus of the Protestant Bible in seventeen English
translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta
for the New Testament.
Looking for every language? This repository is a curated English set,
chosen for spread across translation families and small enough to load whole.
For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses —
see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.jw_myanmar_bible_dataset
📖 JW Myanmar Bible Dataset (New World Translation)
A richly structured, fully aligned dataset of the Myanmar (Burmese) Bible, translated by Jehovah's Witnesses from the New World Translation. This dataset includes 66 books, 1,189 chapters, and 31,078 verses, each with chapter-level URLs and verse-level breakdowns.
✨ Highlights
- 📚 66 Canonical Books (Genesis to Revelation)
- 🧩 1,189 chapters, 31,078 verses (as parsed from the JW.org Myanmar edition)
- 🔗 Includes… See the full description on the dataset page: https://huggingface.co/datasets/freococo/jw_myanmar_bible_dataset.darija_bible
Darija Bible — Moroccan Standard Translation (MSTD)
The full New Testament translated into Moroccan Darija
(الترجمة المغربية القياسية, MSTD).
Moroccan Darija is a low-resource spoken Arabic variety. This dataset is published
here as a research artifact for language modeling, fine-tuning, machine translation
between Darija and other languages, and evaluation of multilingual / Arabic-dialect
models.
⚠️ Copyright notice — please read before using
This dataset reproduces… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/darija_bible.Bible
nwtsty: New World Translation of the Holy Scriptures (Study Edition)
nwt: New World Translation of the Holy Scriptures (2013 Revision)
Rbi8: New World Translation of the Holy Scriptures (1984 Edition)
sbi: Synodal translation
Source: https://www.jw.org
