datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KJV_Bible_Embedding_Collectionkjvnet-kjvlanguages:
en
task_categories:
translation
licenses:
unknown
Dataset Card for [Needs More Information]
Dataset Summary
This is a dataset made up of two Bible translations-- NET and KJV.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
English
Dataset Structure
Data Instances
[Needs More Information]
Data Fields
[Needs More Information]
Data Splits
[Needs More Information]… See the full description on the dataset page: https://huggingface.co/datasets/swcrazyfan/net-kjv.kjvchaptersKJV_PericopesThis is the KJV in JSON, grouped by pericope.
KJVPericopeTopics_bertopicWith bertopic (https://maartengr.github.io/BERTopic/), I ran a dataset of pericopes which covers the entire Bible.
The pericope became the topic, and under each heading, 3 verses were selected as representative.
A useful feature of this is that the representative verses are guaranteed to come from the section of Scripture
that's connected with the pericope, which gives much better quality than if they were, for instance, semantically
chosen from a vector database of the entire Bible.
kjv-strongssKJV-LLM-pretraining.jsonlCleaned_KJV_Bible_for_LLMs
