datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biblical-names-by-en-wikipedia
names generator
from gist description :: https://gist.github.com/Sarverott/729b53dcb1d8695017004294247460fb
example of use package for handling articles on, here using biblical names
( https://en.wikipedia.org/wiki/List_of_biblical_names ) to prepare dataset for hugging face
( https://huggingface.co/datasets/Apokryf/biblical-names-by-en-wikipedia ) to make some
generative human-readable namespace domains for servers.
searching for full content files… See the full description on the dataset page: https://huggingface.co/datasets/Apokryf/biblical-names-by-en-wikipedia.conon-biblical-sft-am-en
📖 Holy-AI-SFT: Biblical Amharic-English SFT Dataset 📖
A high-quality, perfectly balanced Supervised Fine-Tuning (SFT) dataset derived from the canonical bilingual Bible dataset.
🌟 Dataset Overview
This dataset is a specialized collection designed to bridge the gap between ancient scripture and modern AI. It focuses on Supervised Fine-Tuning (SFT) for models requiring deep understanding of biblical context, cross-lingual retrieval, and precise translation… See the full description on the dataset page: https://huggingface.co/datasets/Nexuss0781/conon-biblical-sft-am-en.biblical-tutor-dataset-chirho
Biblical Language Tutor Dataset
For God so loved the world that he gave his only begotten Son, that whoever believes in him should not perish but have eternal life. - John 3:16
Description
Training data for the Biblical Language Tutor pipeline: morphological parsing and interlinear glossing of biblical Hebrew and Greek. Derived from the Macula Hebrew and Greek treebanks (Clear-Bible).
Dataset Structure
Parser Dataset (~200K examples)
JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/LoveJesus/biblical-tutor-dataset-chirho.conon-biblical-am-en
Canon Biblical Amharic-English Dataset (Nexuss0781)
This dataset provides a comprehensive, unified parallel corpus of the Holy Bible in Amharic and English. It is specifically designed for Natural Language Processing (NLP) tasks, including machine translation, cross-lingual information retrieval, and biblical linguistic studies. The dataset aligns the Amharic text with the English New American Standard Bible (NASB) at the verse level.
📖 Dataset Overview
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nexuss0781/conon-biblical-am-en.filipino-tts-biblical
Filipino TTS Dataset (Biblical)
Dataset Description
High-quality Filipino (Tagalog) text-to-speech dataset extracted from biblical audio narration.
Dataset Statistics
Total clips: 630
Total duration: 0.26 hours
Clip duration: 1-2 seconds
Sample rate: 22,050 Hz
Channels: Mono
Language: Filipino (Tagalog)
Source: Ang Dating Biblia audio recordings
Data Format
Each entry contains:
audio_filepath: Path to WAV file
text: Transcribed Filipino text… See the full description on the dataset page: https://huggingface.co/datasets/RidheshBhati/filipino-tts-biblical.even_speech_biblicalThis dataset consists of audiofiles with a speech in Even language.
The correspondence between text and audio is in the table metadata.csv.
The data was collected from religious texts written down by Institute for Bible Translation
Һөвки Дукундукун укчэнэкэл. Institute for Bible Translation, Moscow, 2018.
Притчал. Institute for Bible Translation, Moscow, 2019.
Sourse
Dialect
Total length (min)
Religious texts
Lamunkhin
67.96
Another dataset of Even speech:
field records of… See the full description on the dataset page: https://huggingface.co/datasets/tbkazakova/even_speech_biblical.Nogai-Russian-SFT-Biblical-v1
Nogai-Russian SFT Biblical Corpus v1 (superseded — use v2)
Use Nogai-Russian-SFT-Biblical-v2 instead.
This version is kept unchanged because the published SFT adapter was trained on it. Its splits are not suitable for evaluation (see below).
Russian↔Nogai translation instructions in ChatML format, from human Bible translations by the Institute for Bible Translation (IBT). Used for Phase 2 SFT of NogaiLLM.
What is in it (measured September 2026)
Rows… See the full description on the dataset page: https://huggingface.co/datasets/ansarzeinulla/Nogai-Russian-SFT-Biblical-v1.biblical-ner-dataset-chirho
Biblical NER Dataset (Chirho)
BIO-tagged Named Entity Recognition dataset built from the King James Version (KJV)
Bible text with entity annotations from STEPBible TIPNR data and curated divine name lists.
Format
JSONL with fields:
tokens_chirho: List of word tokens
ner_tags_chirho: List of BIO tags (one per token)
reference_chirho: Bible verse reference
Entity Types
PERSON: Biblical persons (Moses, David, Paul, etc.)
DIVINE: Names and titles of God (God… See the full description on the dataset page: https://huggingface.co/datasets/LoveJesus/biblical-ner-dataset-chirho.biblical-topical-dataset-chirhobiblical-embedding-dataset-chirhofilipino-tts-biblical-fullbiblical-variant-dataset-chirhobiblical-systems-methodadaption-igbo-biblical-texts
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-igbo_biblical_texts
This dataset contains a collection of text completions in the Igbo language, primarily focusing on biblical narratives, religious teachings, and scriptural quotes. The samples include references to figures like Jesus and Abraham, as well as descriptions of miracles and moral exhortations found in Christian theology. The content is formatted as standalone sentences… See the full description on the dataset page: https://huggingface.co/datasets/Khaycee/adaption-igbo-biblical-texts.Nogai-Russian-SFT-Biblical-v2
Nogai-Russian SFT Biblical Corpus v2
Russian↔Nogai translation instructions (ChatML) from human Bible translations by the Institute for Bible Translation (IBT). This is a clean rebuild of v1 with splits that can be used for evaluation. Built with build_sft_clean.py (seed 42).
Splits
Split
Pairs
Rows (2 directions per pair)
train
507
1,014
validation
60
120
test
58
116
How it was built from v1
4,310 v1 rows reduce to 650 unique… See the full description on the dataset page: https://huggingface.co/datasets/ansarzeinulla/Nogai-Russian-SFT-Biblical-v2.
