datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Latin-PD
🇲🇪 Latin Public Domain Books (Latin) 🇲🇪
Latin-Public Domain or Latin-PD is a large collection aiming to aggregate all Latin monographies and periodicals in the public domain. As of June 2024, it is the largest Latin open corpus.
Dataset summary
The collection contains 16,521,454,086 words (159,070 titles) recovered from multiple sources, including the Internet Archive and various European national libraries and cultural heritage institutions (BDH, BNF). Each parquet… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Latin-PD.LatinFontsSVGs
SVG Font Dataset
Overview
We present a curated dataset of vector font glyphs stored as SVG files, designed for research in generative modeling, structured vector graphics and typography synthesis.
The dataset was created for the development and evaluation of our paper:
DesigNet: Learning to Draw Vector Graphics as Designers Do
Related Resources
Paper (arXiv) : https://arxiv.org/abs/2604.06494
Code: https://github.com/TomasGuija/DesigNet… See the full description on the dataset page: https://huggingface.co/datasets/TomasGuija/LatinFontsSVGs.latin-classical-intertextuality-labels
Latin Jerome Intertextuality Labels
This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.LatinEpicIntertextualityRetrieval
Latin Epic Intertextuality Retrieval
BEIR-style Retrieval benchmark for Latin epic intertextuality under PoetryMTEB.
Given a passage from Valerius Flaccus, Argonautica Book 1, retrieve the corresponding verse line(s) in Vergil (Aeneid), Lucan (Bellum Civile), Ovid (Metamorphoses), or Statius (Thebaid) that traditional scholarship identifies as parallels.
Gold parallels: Burns et al., NAACL-HLT 2021 (paper; repo), 945 curated pairs.
Dataset Card
Item… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/LatinEpicIntertextualityRetrieval.Latin-CC-170M
Latin-CC-170M
Reupload of the Corpus Corporum as Parquet, originally from
Kaggle,
and reuploaded as CSV by
Fece228/latin-literature-dataset-170M.
License
These works are public domain.
Original README
This is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin up to 19th century, which includes:
Classical Latin: works of Caesar, Cicero and many more
Medieval Latin: a… See the full description on the dataset page: https://huggingface.co/datasets/AncientLanguages/Latin-CC-170M.Latin-OCR-Artifacts
Latin sentences sourced from The Latin Library, converted to images, were subsequently degraded via OCRODEG.
OCR (via Kraken and Tesseract) transcriptions were generated. The dataset was then augmented with several synthetic noise patterns in order to emulate the more severe corruption found in many older digitizations.
If you use this in your work, please cite:
@misc{mccarthy2025LACOROCR,
author = {McCarthy, A. M.},
title = {{Latin OCR Artifacts}},
year = {2025}… See the full description on the dataset page: https://huggingface.co/datasets/aimgo/Latin-OCR-Artifacts.Pleias-Latin-PD-cleanedlatin-german-parallel
📜 Latin-German Textcorpus
This dataset consists of 406,011 Latin-German parallel sentences (sentence pairs).
Each entry contains a Latin sentence and its corresponding German translation.
The sentence pairs were collected and processed from various websites and online sources.
📄 Dataset Schema
The dataset contains the following columns:
id: A unique identifier for each entry.
latin: The sentence in Latin.
german: The German translation.
source: The origin or reference… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/latin-german-parallel.maldivian-latin-script-corpus
Maldivian Latin-Script Corpus
A growing collection of Latin-script text from Maldivian online communities, containing Romanized Dhivehi, English, and code-mixed writing. Data is collected from multiple sources and tagged by origin.
Why this dataset is unique
Maldivians commonly write Dhivehi phonetically using Latin script rather than switching to the Thaana keyboard. This produces text like:
"varah reethi vaahaka eh" → ވަރަށް ރީތި ވާހަކައެއް (very nice story)
"maa salhi… See the full description on the dataset page: https://huggingface.co/datasets/d3b4g/maldivian-latin-script-corpus.latin_vulgate_la
Biblia Sacra Vulgata (4th century)
Description
The Biblia Sacra Vulgata is Jerome's Latin translation of the Bible, completed around 405 AD. Commissioned by Pope Damasus I in 382, Jerome translated directly from the Hebrew (Old Testament) and Greek (New Testament) texts. The Vulgate became the standard Latin Bible of the Western Church for over 1,500 years. This edition represents the pre-Clementine Vulgate text, distinct from the Clementine Vulgate (1592) already… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/latin_vulgate_la.
