datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
catholic-resources
Vietnamese Catholic resources by v-bible
Data Structure
calendar: Generated Liturgical calendars using
v-bible/js-sdk.
misc/proper-names.json: Name translation from
ktcgkpv.org, generated by
v-bible/bible-scraper.
liturgical: Liturgical data from
The Lectionary for Mass (1998/2002 USA Edition),
compiled by Felix Just, S.J., Ph.D., and generated by
v-bible/bible-scraper.
books/bible: Generated Bible markdown data.
books/catechism-books: Official catechism… See the full description on the dataset page: https://huggingface.co/datasets/v-bible/catholic-resources.kjv-biblebible-embeddings
Bible Embeddings
A comprehensive tool for generating and evaluating Bible verse embeddings using various state-of-the-art embedding models. This project supports both commercial APIs (OpenAI, Google Gemini, Voyage AI) and open-source models (HuggingFace sentence-transformers) for semantic search across biblical texts.
Setup
This project is managed with uv. Make sure you have uv installed, then set up the project:
# Install dependencies
uv sync
# Install specific… See the full description on the dataset page: https://huggingface.co/datasets/LetsChurch/bible-embeddings.open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.open-bible-resourcesopen-bible-speech-african
Open Bible Resources — African Languages
Spoken-audio Bible recordings aligned to verse-level text for 19 African languages —
roughly 1,741 hours of audio across ~552,907 audio–text pairs (~357 GB).
This dataset is the African-language subset of
davidguzmanr/open-bible-resources,
re-hosted here by AfriSpeech to make the African
languages easy to find and use on their own. The audio and text are unchanged from the
source; only the non-African configurations have been removed. All… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/open-bible-speech-african.biblenlp-corpus-mmtebThis dataset pre-computes all English-centric directions from bible-nlp/biblenlp-corpus, and as a result loading is significantly faster.
Loading example:
>>> from datasets import load_dataset
>>> dataset = load_dataset("davidstap/biblenlp-corpus-mmteb", "eng-arb", trust_remote_code=True)
>>> dataset
DatasetDict({
train: Dataset({
features: ['eng', 'arb'],
num_rows: 28723
})
validation: Dataset({
features: ['eng', 'arb'],
num_rows: 1578
})… See the full description on the dataset page: https://huggingface.co/datasets/davidstap/biblenlp-corpus-mmteb.biblenlp-corpus-mmteb
BibleNLPBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
Partial Bible translations in 829 languages, aligned by verse.
Task category
t2t
Domains
Religious, Written
Reference
https://arxiv.org/abs/2304.09919
Source datasets:
davidstap/biblenlp-corpus-mmteb
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("BibleNLPBitextMining")
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biblenlp-corpus-mmteb.asante-twi-bible-speech-text
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
sign-bibles
bible-nlp/sign-bibles
This dataset is still being generated and currently includes only test files
This dataset contains sign language videos from the Digital Bible Library (DBL), processed for machine learning applications. The dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0).
Dataset Structure
Each sample contains:
["mp4"] the original video
["json"] Metadata, including bible reference, copyright information… See the full description on the dataset page: https://huggingface.co/datasets/bible-nlp/sign-bibles.biblenlp-corpusThis dataset pre-computes all English-centric directions from bible-nlp/biblenlp-corpus, and as a result loading is significantly faster.
Loading example:
>>> from datasets import load_dataset
>>> dataset = load_dataset("davidstap/biblenlp-corpus-mmteb", "eng-arb", trust_remote_code=True)
>>> dataset
DatasetDict({
train: Dataset({
features: ['eng', 'arb'],
num_rows: 28723
})
validation: Dataset({
features: ['eng', 'arb'],
num_rows: 1578
})… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biblenlp-corpus.cameroon_bibles
Cameroon Bibles — verse-aligned scripture corpus
The text corpus behind Lingo / NativeAI: verse-aligned scripture
across 60 Cameroonian languages (64 translation versions). Scripture is one of
the few sources of sentence-aligned parallel text for these low-resource languages — the
aligned backbone of our corpus (see the research log).
Layout
<Language>/<BOOK>.<chapter>.txt e.g. Ngi/MAT.2.txt
Each file is one chapter; lines are verse-numbered, alignable across… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/cameroon_bibles.ewe-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
48775 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.suvartha-bible-audio-irv
Suvartha Bible audio — CC BY-SA 4.0
Chapter-by-chapter readings of the Bible, one MP3 per chapter, as played in the Bible reader at https://suvartha.in.
Folder
Language
Credit
Source
Licence
hi-irv
Hindi
Indian Revised Version — text and audio © Bridge Connectivity Solutions / Davar Partners International, CC BY-SA 4.0 (via Snow Mountain dataset)
source
CC BY-SA 4.0
kn-irv
Kannada
Indian Revised Version — text and audio © Bridge Connectivity Solutions / Davar… See the full description on the dataset page: https://huggingface.co/datasets/roycvn/suvartha-bible-audio-irv.The-Bible-KJVbiblenlp-corpus
Dataset Card for BibleNLP Corpus
Dataset Summary
Partial and complete Bible translations in 833 languages, aligned by verse.
Languages
aai, aak, aau, aaz, abt, abx, aby, acf, acr, acu, adz, aer, aey, agd, agg, agm, agn, agr, agt, agu, aia, aii, aka, ake, alp, alq, als, aly, ame, amf, amk, amm, amn, amo, amp, amr, amu, amx, anh, anv, aoi, aoj, aom, aon, apb, ape, apn, apr, apu, apw, apz, arb, are, arl, arn, arp, asm, aso, ata, atb, atd, atg, att, auc, aui, auy… See the full description on the dataset page: https://huggingface.co/datasets/bible-nlp/biblenlp-corpus.bible_king_james_version_en
King James Version (1611)
Description
The most influential English Bible translation in history, commissioned by King James I of England and first published in 1611. The translation was prepared by 47 scholars organized into six committees, working from the original Hebrew, Aramaic, and Greek texts, as well as consulting earlier English translations (Tyndale, Coverdale, Geneva Bible) and the Latin Vulgate. The KJV is renowned for the majesty of its prose and its… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/bible_king_james_version_en.bible_tts_hausa
Dataset Card for BibleTTS Hausa
Dataset Summary
BibleTTS is a large high-quality open Text-to-Speech dataset with up to 80 hours of single speaker, studio quality 48kHz recordings.
This is a Hausa part of the dataset. Aligned hours: 86.6, aligned verses: 40,603.
Languages
Hausa
Dataset Structure
Data Fields
audio: audio path
sentence: transcription of the audio
locale: always set to ha
book: 3-char book encoding
verse: verse id… See the full description on the dataset page: https://huggingface.co/datasets/vpetukhov/bible_tts_hausa.dagbani-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.BibleMMSThe Dataset associated with the Paper "Meta Learning Text-to-Speech Synthesis in over 7000 Languages" by Florian Lux, Sarina Meyer, Lyonel Behringer, Frank Zalkow, Phat Do, Matt Coler, Emanuël A. P. Habets and Ngoc Thang Vu (Interspeech 2024).
We generate 2000 spoken utterances per language using the subsets of the eBible dataset [1] that are under free licenses as the text input to the MMS TTS models [2].
The languages associated with the following ISO-639-3 codes are represented in this… See the full description on the dataset page: https://huggingface.co/datasets/Flux9665/BibleMMS.bible-text-fixing-v0zo-bible
zo-bible
A sentence-aligned parallel Bible corpus covering 8 closely related Zo speech varieties and English across 10 translation versions (30,974 canonical verses).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr), Mizo (lus), Paite (pck), Vaiphei (vap), Thadou (tcz), Gangte (gnb), Zou (zom), English (eng)
Family: Zo Languages
Volume: 30,974 verse anchors across 66 canonical books (10 translation editions)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/zo-bible.bible-sphere-statsbiblenlp-corpus
BibleNLP Corpus
This is a conversion of BibleNLP corpus to the Parquet format,
Dataset Summary
The dataset contains partial and complete Bible translations in 835 languages, aligned by verse. Each language is stored as a separate Parquet file (eng.parquet, fra.parquet, …).
This format is derived from the eBible corpus corpus.json and is intended for fast columnar loading with Hugging Face datasets, Polars, Pandas, or DuckDB.
Languages
835 ISO 639-3… See the full description on the dataset page: https://huggingface.co/datasets/nordpolemil/biblenlp-corpus.kikongo-bible-asr
Kikongo Bible ASR
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/kikongo-bible-asr.bible_paraThis is a multilingual parallel corpus created from translations of the Bible compiled by Christos Christodoulopoulos and Mark Steedman.
102 languages, 5,148 bitexts
total number of files: 107
total number of tokens: 56.43M
total number of sentence fragments: 2.84Mbible-new-testamentshona-bible-bdsc-aligned
Shona Bible Speech Alignment Dataset
Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC
source audio made available by Biblica, Inc. through Open.Bible. This release
contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech
segments covering approximately 75.55 hours.
Dataset summary
Language: Shona (sna)
Speaker: narrator 1
Speaker sex: male
Books: 66
Clips: 31,284
Audio: approximately 75.55 hours
Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.biblesMultilingual Biblesbiblelm
BibleLM Dataset
A high-performance, stateless Bible dataset optimized for edge-first RAG (Retrieval-Augmented Generation).
This dataset contains the processed Bible text, morphological data, and search indices used by the BibleLM project.
📚 What's inside?
Combined Bible Index: Cleaned and tokenized text for BSB (Berean Standard Bible), KJV, WEB, and ASV.
Search Engine State: Pre-computed BM25 term frequencies (bm25-state.json) allowing for <10ms search engine… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/biblelm.
