catholic
Datasets
All datasets matching “catholic”catholic-resources
Vietnamese Catholic resources by v-bible
Data Structure
calendar: Generated Liturgical calendars using
v-bible/js-sdk.
misc/proper-names.json: Name translation from
ktcgkpv.org, generated by
v-bible/bible-scraper.
liturgical: Liturgical data from
The Lectionary for Mass (1998/2002 USA Edition),
compiled by Felix Just, S.J., Ph.D., and generated by
v-bible/bible-scraper.
books/bible: Generated Bible markdown data.
books/catechism-books: Official catechism… See the full description on the dataset page: https://huggingface.co/datasets/v-bible/catholic-resources.catholiccorpus
CatholicCorpus
An open-access, NLP-ready corpus of Catholic texts spanning 2,000 years of the Catholic intellectual tradition — from the Church Fathers to the 20th century.
67,772 content files | 16 collections | 35.9 GB | 2,000 years of coverage
What's In the Corpus
#
Collection
Content Files
Size
Format
01
Git Repos (CSEL, Aquinas Opera Omnia, Septuagint, Byzantine Text, eBible)
4,668
3.4 GB
TEI XML, TXT
02
Corpus Corporum (Patrologia Latina + 29… See the full description on the dataset page: https://huggingface.co/datasets/CatholicCorpus/catholiccorpus.catholiccorpus-text
CatholicCorpus — Extracted Text
Pre-extracted plain text from the CatholicCorpus — 2,000 years of the Catholic intellectual tradition, ready for NLP, RAG, and digital humanities.
This dataset contains 47,407 plain text files (5.7 GB, 2.64 billion GPT-2 tokens) extracted from the raw source corpus (PDF, EPUB, TEI XML, HTML). If you need the original source formats, see the raw corpus.
Quick Start
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/CatholicCorpus/catholiccorpus-text.Catholic_Datasetsarch-unintel-passers-lpi-260903T1640-catholicism-k30-rowsarch-unintel-passers-lpi-260903T1640-catholicism-k30-adapters
