datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sermon-pair
Sermon Transcript Dataset
Description
This dataset is a curated collection of approximately 960,000 rows of sermon transcripts, obtained from sermons uploaded on YouTube. Each sermon was transcribed, split into smaller passage-like segments, and then manually labeled over a span of 4 weeks. The labeling process extracted the last Bible passage mentioned in each segment.
Purpose
The dataset was created to facilitate the development of a plugin for EasyWorship… See the full description on the dataset page: https://huggingface.co/datasets/odunola/sermon-pair.unitarian-universalist-sermons
6900 transcripts
44 churches
timeframe: 2010-2022
Denomination: Unitarian Universalist, USA
Dataset structure
church (church name or website)
source (mp3 file)
text
sentences (count)
errors (number of sentences skipped because could not understand audio, or just long pauses skipped)
duration (in seconds)
Dataset creation
see notebook in files
check-if-sermon
