datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
The-Bible-KJVbiblecorpuscsvBiblebible-caucasus
Bible Translations in Languages of the Caucasus
Verse-aligned Bible translations in 22 languages of the Caucasus (plus Russian and
English as pivots). Most are published by the
Institute for Bible Translation (IBT, Moscow); the Udi edition
comes from Translation Services International; the two
English editions (WEB, KJVA) are public-domain reference translations from the same IBT
server (see Provenance). One config per language/edition, 400,286 verse rows total.
Every row… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/bible-caucasus.restructured-biblecorpusI do not hold the copyright to this dataset; I merely restructured it to have the same structure as other datasets (that we are researching) to facilitate future coding and analysis. I refer to this link for the raw dataset.
ro-paraphrase-bible
Dataset Card for "Romanian Bible Paraphrase Corpus"
Dataset Description
Homepage: https://github.com/AndyTheFactory/ro-paraphrase-bible
Repository: https://github.com/AndyTheFactory/ro-paraphrase-bible
Point of Contact: Andrei Paraschiv
Dataset Summary
A paraphprase corpus created from 10 different Romanian language Bible versions. Since the Bible has all paragraphs uniquely numbered an alignment between two
versions is straighforward.
We compiled a… See the full description on the dataset page: https://huggingface.co/datasets/andyP/ro-paraphrase-bible.bible-reference-sentence-pairKonkani_New_testament_bibleEhn-bible-bbc-gpt3.5
Dataset Card for Ehn-Bible-BBC-GPT3.5
Dataset Summary
This dataset card contains parallel Nigerian Pidgin and English sentences split into three files, namely: train.csv, valid.csv and test.csv.
The original data was split in the ratio of 8:1:1 to obtain these files.
Supported Tasks and Leaderboards
Language Translation
Language Identification
Languages
English
Nigerian Pidgin
Dataset Structure
Data Instances
Data… See the full description on the dataset page: https://huggingface.co/datasets/NITHUB-AI/Ehn-bible-bbc-gpt3.5.jw_myanmar_bible_dataset
📖 JW Myanmar Bible Dataset (New World Translation)
A richly structured, fully aligned dataset of the Myanmar (Burmese) Bible, translated by Jehovah's Witnesses from the New World Translation. This dataset includes 66 books, 1,189 chapters, and 31,078 verses, each with chapter-level URLs and verse-level breakdowns.
✨ Highlights
- 📚 66 Canonical Books (Genesis to Revelation)
- 🧩 1,189 chapters, 31,078 verses (as parsed from the JW.org Myanmar edition)
- 🔗 Includes… See the full description on the dataset page: https://huggingface.co/datasets/freococo/jw_myanmar_bible_dataset.bible-lezghian-russianBible
Full Bible Chapter wise - Tamil
Web Scrapped from https://bible.catholicgallery.org/ecu-tamil/
english-ceb-bible-prompt
LLM Benchmark for English-Cebuano Translation
This dataset contains parallel sentences of English and Cebuano extracted from the Bible corpus available at https://github.com/christos-c/bible-corpus. The dataset is formatted for use in training machine translation models, particularly with the Transformers library from Hugging Face.
Usage
This dataset can be used to evaluate the performance of Large Language Models for English-Cebuano machine translation using libraries… See the full description on the dataset page: https://huggingface.co/datasets/eemberda/english-ceb-bible-prompt.bible-passage-transcriptionparallel-catholic-bible-versions
Parallel Catholic Bible Versions
Aligned verses from all 73 books of Catholic Bible in three versions: the Latin Vulgate (vulgate), the Catholic Public Domain Version (cpdv), and the Douay-Rheims Challoner Revision (drc). Includes 2,450 translation commentary notes from the Latin English Study Bible by Ronald L. Conte Jr.
Sources
Latin-English Study Bible with notes scraped by aseemsavio
Douay-Rheims via scrollmapper
Text has been cleaned to remove HTML tags (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/jam963/parallel-catholic-bible-versions.Ehn-bible
Dataset Card for Ehn-Bible-BBC-GPT3.5
Dataset Summary
This dataset card contains parallel Nigerian Pidgin and English sentences split into three files, namely: train.csv, valid.csv and test.csv.
The original data was split in the ratio of 8:1:1 to obtain these files.
Supported Tasks and Leaderboards
Language Translation
Language Identification
Languages
English
Nigerian Pidgin
Dataset Structure
Data Instances
english… See the full description on the dataset page: https://huggingface.co/datasets/NITHUB-AI/Ehn-bible.twi_bible_v1
Twi Text-to-Speech
tachiwin_biblesFunny-Windows-Errors-Windows-Biblesewe_bible_v1
Ewe bible for Text-to-Speech
bible-reference-sentence-pairbible-nl-enenglish-ceb-bible
English-Cebuano Bible Translation Dataset
This dataset contains parallel sentences of English and Cebuano extracted from the Bible corpus available at https://github.com/christos-c/bible-corpus. The dataset is formatted for use in training machine translation models, particularly with the Transformers library from Hugging Face.
Usage
This dataset can be used to fine-tune pre-trained language models for English-Cebuano machine translation using libraries like Transformers.… See the full description on the dataset page: https://huggingface.co/datasets/eemberda/english-ceb-bible.eng_igl_bibleraw-indigenous-bible
Scriptures Translation Project
Welcome to the Scriptures Translation Project! This repository contains information about translations of religious scriptures in various languages. Below is a table summarizing two translations available in Guajajara and Guarani languages for Brazil.
Translations
Description
Abbreviation
Comments
Version
VersionDate
PublishDate
RightToLeft
OT
NT
Strong
The Scriptures in Guajajara of Brazil. [gub]
Guajajara
The Scriptures in… See the full description on the dataset page: https://huggingface.co/datasets/tiagoblima/raw-indigenous-bible.darija_bible
Darija Bible — Moroccan Standard Translation (MSTD)
The full New Testament translated into Moroccan Darija
(الترجمة المغربية القياسية, MSTD).
Moroccan Darija is a low-resource spoken Arabic variety. This dataset is published
here as a research artifact for language modeling, fine-tuning, machine translation
between Darija and other languages, and evaluation of multilingual / Arabic-dialect
models.
⚠️ Copyright notice — please read before using
This dataset reproduces… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/darija_bible.bible-nt-datasetcorpus-paite-english-bibleenglish-tumbuka-bibleeng_edo_bible
