datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biblenlp-corpus-mmtebThis dataset pre-computes all English-centric directions from bible-nlp/biblenlp-corpus, and as a result loading is significantly faster.
Loading example:
>>> from datasets import load_dataset
>>> dataset = load_dataset("davidstap/biblenlp-corpus-mmteb", "eng-arb", trust_remote_code=True)
>>> dataset
DatasetDict({
train: Dataset({
features: ['eng', 'arb'],
num_rows: 28723
})
validation: Dataset({
features: ['eng', 'arb'],
num_rows: 1578
})… See the full description on the dataset page: https://huggingface.co/datasets/davidstap/biblenlp-corpus-mmteb.biblenlp-corpus-mmteb
BibleNLPBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
Partial Bible translations in 829 languages, aligned by verse.
Task category
t2t
Domains
Religious, Written
Reference
https://arxiv.org/abs/2304.09919
Source datasets:
davidstap/biblenlp-corpus-mmteb
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("BibleNLPBitextMining")
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biblenlp-corpus-mmteb.biblenlp-corpusThis dataset pre-computes all English-centric directions from bible-nlp/biblenlp-corpus, and as a result loading is significantly faster.
Loading example:
>>> from datasets import load_dataset
>>> dataset = load_dataset("davidstap/biblenlp-corpus-mmteb", "eng-arb", trust_remote_code=True)
>>> dataset
DatasetDict({
train: Dataset({
features: ['eng', 'arb'],
num_rows: 28723
})
validation: Dataset({
features: ['eng', 'arb'],
num_rows: 1578
})… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biblenlp-corpus.bible-sphere-statsbiblelm
BibleLM Dataset
A high-performance, stateless Bible dataset optimized for edge-first RAG (Retrieval-Augmented Generation).
This dataset contains the processed Bible text, morphological data, and search indices used by the BibleLM project.
📚 What's inside?
Combined Bible Index: Cleaned and tokenized text for BSB (Berean Standard Bible), KJV, WEB, and ASV.
Search Engine State: Pre-computed BM25 term frequencies (bm25-state.json) allowing for <10ms search engine… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/biblelm.KJV_Bible_Embedding_Collectionbible-reference
Bible Reference Corpus
Thirteen aligned reference datasets for study of the biblical text: Greek and
Hebrew lexicons keyed to Strong's numbers, an interlinear word map, the critical
apparatus of eight Greek editions, cross-reference and topical indexes, and
geolocated places.
Published by SermonIndex.
Everything in this repository is public domain or CC BY 4.0. Sources with
share-alike terms are kept in a separate repository,
sermonindex/bible-reference-sa,
so that a share-alike… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-reference.BibleDictionaries
JSON for:
Easton's Bible Dictionary
Smith's Bible Dictionary
Hitchcock's Bible Names Dictionary
Torry's Topical Handbook
bible
The Bible in 1,004 Languages
14,497,397 verses across 1,253 translations in 1,004 languages, every verse
keyed to the same chapter-and-verse address so that any two languages can be
aligned by joining on book, chapter and verse.
The Bible is the most widely translated text in existence, and for several
hundred of the languages here it is the largest — sometimes the only —
substantial digitised text. That makes this corpus unusually useful for
low-resource machine translation… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible.bible-reference-sa
Bible Reference Corpus — Share-Alike Tier
The part of the SermonIndex Bible reference corpus that derives from
CC BY-SA 4.0 sources, kept in its own repository so that the share-alike
obligation is explicit and does not spread to the permissively licensed
material.
The main corpus is
sermonindex/bible-reference
(CC BY 4.0). Join on strongs / key.
If you use this repository, your derivative work must also be share-alike.
If that does not suit you, use the CC BY 4.0 repository… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-reference-sa.bible-datasets-ptbible-parallel-english
Parallel Bible — English Translations and Ancient Versions
A verse-aligned parallel corpus of the Protestant Bible in seventeen English
translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta
for the New Testament.
Looking for every language? This repository is a curated English set,
chosen for spread across translation families and small enough to load whole.
For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses —
see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.bibletts-asante-twi-repaired
BibleTTS Asante Twi — Repaired Transcripts
The Asante Twi transcripts released with BibleTTS have had the
characters ɛ (U+025B) and ɔ (U+0254) stripped out. This dataset restores them.
Audio is not included. This is a drop-in replacement for the .txt files that ship with the
BibleTTS Asante Twi package, matched by clip ID.
The problem
Both are Twi vowels, and both are required by the orthography. Measured across the released
Asante Twi transcripts:
Character… See the full description on the dataset page: https://huggingface.co/datasets/danieldzikunuofmarvel/bibletts-asante-twi-repaired.LeroyDyer___Spydaz_Web_AI_BIBLE_002-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_BIBLE_002
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_BIBLE_002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_BIBLE_002-details.Baptist-Christian-Bible-Expert
Updated dataset and updated guide!
Comprehensive Guide for QLoRA Fine Tuning
1. Initial Guide Setup:
You can make this cut & paste easy by finding and replacing the following variables in the guide. Copy over the whole thing including brackets.
Point to your local files.
[local_pc_path_to_config_and_data]
[config.yml]
[dataset.jsonl]
Pick a name.
[runpod_model_folder_name]
SSH connection to runpod.
[serverIP]
[sshPort]
How will you upload your model will go on HF?… See the full description on the dataset page: https://huggingface.co/datasets/sleepdeprived3/Baptist-Christian-Bible-Expert.sango-french-bible-parallel
SFPC: Sango-French Parallel Corpus
The first quality-filtered, verse-aligned Sango-French parallel corpus, constructed for neural machine translation research. This dataset directly addresses the "Sango Problem" identified by Meta's NLLB-200 project — the failure of cross-lingual transfer for a linguistically isolated Creole language.
Associated resources:
Model: alaminerca/nllb-sango-french
Demo: Sango-French Translator
Paper: SangoNMT: Parameter-Efficient Domain Adaptation of… See the full description on the dataset page: https://huggingface.co/datasets/alaminerca/sango-french-bible-parallel.biblechaptersReformed-Baptist-1689-Bible-Expert-v7.0cherokee-english-bible-7.96k
Cherokee-English Bible Dataset (8k)
Overview
The Cherokee-English Bible Dataset is a specialized collection of 8,000 entries, each containing a verse from the Cherokee Bible along with its English translation. This dataset is a valuable resource for linguistic scholars, theologians, and developers working on language processing tools that require a deep understanding of both the Cherokee and English languages, particularly within the context of religious texts.… See the full description on the dataset page: https://huggingface.co/datasets/wang4067/cherokee-english-bible-7.96k.Bible-responses-dataset-gotquestions
Theology Question-Answer Dataset
Description
This dataset contains structured, human-generated content focused on theology, primarily sourced from the website GotQuestions. Each entry is formatted as a question (prompt) and a corresponding answer (response). The dataset is provided in JSON format and is intended for fine-tuning AI models, though it can be used for other purposes as well.
The structure of the dataset is as follows:
{
"prompt": "What does it mean to… See the full description on the dataset page: https://huggingface.co/datasets/vericudebuget/Bible-responses-dataset-gotquestions.bible-dpoDatadudeDev/Bible with an added rejected column made using Phi3-q4.
bible_topicsThese are topics with their verse reference. Some of them have cross-references, and some of the cross-references have been voted on.
topic_scores.json and topic_votes.json are both from openbible.info, retrieved November 1, 2023.
biblejsonReformed-Christian-Bible-Expert
QLoRA Fine-Tuning
1. Runpod Setup
Template: runpod/pytorch:2.2.0-py3.10-cuda12.1.1-devel-ubuntu22.04
Expose SSH port (TCP): YOUR_PORT
IP: YOUR_IP
2. Local Machine Preparation
Generate SSH key
3. SSH Connection
SSH over exposed TCP
Connect to your pod using SSH over a direct TCP connection. (Supports SCP & SFTP)
4. Server Configuration
# Update system
apt update && apt upgrade -y
apt install -y git-lfs tmux htop libopenmpi-dev
# Create workspace
mkdir -p… See the full description on the dataset page: https://huggingface.co/datasets/sleepdeprived3/Reformed-Christian-Bible-Expert.biblechatavos-dark-fantasy-lore-bible
THE DARKNESS STEALS THE LIGHT: Authoritative Lore Bible
Author: Mark Kirkbride
Word Count: ~120000
Canon Status: Master Node v3.0
Raw Full Book Text as Nested Json the-darkness-steals-the-light_Full_Text_v2026 // https://huggingface.co/datasets/TheElim/avos-dark-fantasy-lore-bible/raw/main/the-darkness-steals-the-light_Full_Text_v2026.json
Text Jsonl for Book Training https://huggingface.co/datasets/TheElim/avos-dark-fantasy-lore-bible/raw/main/TDSTL_train.jsonl
I. AI… See the full description on the dataset page: https://huggingface.co/datasets/TheElim/avos-dark-fantasy-lore-bible.ota-bible-parallel
Ottoman–Turkish–English Parallel New Testament Corpus
This repository is a verse-aligned parallel corpus for Ottoman Turkish (Perso-Arabic original script) with English and modern Turkish reference translations.
It is intended for training and evaluating translation models for
Ottoman Turkish, a low-resource historical language.
Source language: Ottoman Turkish (ota), Perso-Arabic script
Target languages: English (en), modern Turkish (tr)
Unit of alignment: a single New… See the full description on the dataset page: https://huggingface.co/datasets/enesyila/ota-bible-parallel.gemini-bavarian-bible
Gemini-powered Bavarian Bible
This datasets hosts a translated version of the Public Domain Luther Bible from 1912 from ebible.org. The recently released Gemini 2.5 Pro model was used to translate the German version of the Luther bible into Bavarian.
Dataset Format
Here's an example of the JSON format, that is used for this dataset:
{
"id": "deu1912_002_GEN_01",
"text": "As 1. Buach Mose (Genesis).\n1.\nAm Ofang hot da Herrgott an Himme und d'Erdn gschaffa..."
}
The… See the full description on the dataset page: https://huggingface.co/datasets/bavarian-nlp/gemini-bavarian-bible.bible_dictionary_unifiedThese 4 Bible dictionaries are combined:
-Easton's Bible Dictionary
-Hitchcock's Bible Names Dictionary
-Smith's Bible Dictionary
-Torrey's Topical Textbook
nvidia-nemo-nano-codec-bible
