datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.Latin-PD
🇲🇪 Latin Public Domain Books (Latin) 🇲🇪
Latin-Public Domain or Latin-PD is a large collection aiming to aggregate all Latin monographies and periodicals in the public domain. As of June 2024, it is the largest Latin open corpus.
Dataset summary
The collection contains 16,521,454,086 words (159,070 titles) recovered from multiple sources, including the Internet Archive and various European national libraries and cultural heritage institutions (BDH, BNF). Each parquet… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Latin-PD.LatinFontsSVGs
SVG Font Dataset
Overview
We present a curated dataset of vector font glyphs stored as SVG files, designed for research in generative modeling, structured vector graphics and typography synthesis.
The dataset was created for the development and evaluation of our paper:
DesigNet: Learning to Draw Vector Graphics as Designers Do
Related Resources
Paper (arXiv) : https://arxiv.org/abs/2604.06494
Code: https://github.com/TomasGuija/DesigNet… See the full description on the dataset page: https://huggingface.co/datasets/TomasGuija/LatinFontsSVGs.latin-literature-dataset-170MThis is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin. The dataset is split in two parts: preprocessed with basic cltk tools, ready for work, and raw text data. It must be noted, however, that the latter contains text in Greek, Hebrew, and other languages, with references and contractions
latin-summarizer-dataset
✨ LatinSummarizer Dataset ✨
Note: If Dataset Viewer is not available, see samples of dataset for samples from the dataset.The LatinSummarizer Dataset is a comprehensive collection of Latin texts designed to support natural language processing research for a low-resource language. It provides parallel data for various tasks, including translation (Latin-to-English) and summarization (extractive and abstractive).
This dataset was created for a… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/latin-summarizer-dataset.LatinSummarizer
LatinSummarizer Dataset
Structure
aligned_en_la_data_raw.csv
aligned_en_la_data_cleaned.csv
aligned_en_la_data_cleaned_with_stanza.csv
concat_aligned_data.csv
concat_cleaned.csv
latin_wikipedia_cleaned.csv
latin_wikipedia_raw.csv
latin-literature-dataset-170M_raw_cleaned.csv
latin-literature-dataset-170M_raw_cleaned_chunked.csv
Elsa_aligned/
README.md
Details
aligned_en_la_data_raw.csv
This dataset contains aligned Latin (la) - English (en)… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/LatinSummarizer.arabic-latin-invoices-synthetic
Synthetic Arabic/Latin Invoices — label-first multimodal dataset
A large, perfectly-labeled synthetic invoice dataset for training multimodal
invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin
layouts, and all three numeral glyph systems.
Every sample is generated label-first: a structured invoice record is sampled, then
rendered to pixels via a headless browser, and bounding boxes are read back from the same
DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.uyghur-cv-latinlatin_english_translation
Dataset Card for "latin_english_parallel"
101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation.
For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences.
Each sample is annotated with the index and file (and therefore author/work) that the sample is from. If you find errors… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_translation.cc100-latin
Latin part of cc100 corpus
This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface.
Preprocessing
I undertook the following preprocessing steps:
Removal of all "pseudo-Latin" text ("Lorem ipsum ...").
Use of CLTK for sentence splitting and normalisation.
Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.latin-classical-intertextuality-corpus
Latin Classical Authors Corpus
This dataset contains processed texts from classical Latin authors, serving as a retrieval corpus for intertextuality research. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Cicero and classical Latin literature.
Related Datasets
This corpus is part of the Latin Jerome Intertextuality collection:
Queries: latin-classical-intertextuality-queries - The whole works of… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-corpus.latin-classical-intertextuality-labels
Latin Jerome Intertextuality Labels
This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.latin-classical-intertextuality-queries
Latin Classical Intertextuality Queries
This dataset contains query texts used for finding intertextual relationships with classical Latin authors. It comprises the whole works of Hieronymus (Jerome) and Lactantius, which are searched against a corpus of classical Latin literature. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Lactantius and classical Latin literature.
Related Datasets
This queries… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-queries.WP8-Latin-Embeddings-Indicesmedieval-latinlatin_english_parallel
Dataset Card for "latin_english_parallel"
101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation.
For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences. Additionally, the English translations were both 1. copyrighted and 2. outdated. As such, we decided to modernize and… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_parallel.latin_treebanks_ud_testNOTE: This template for datasheets for ancient language data is based on the proposal by Gebru et al. 2021 https://arxiv.org/abs/1803.09010. The majority of questions is taken from there, but in a slightly rearanged order. However, as some questions are not relevant for historical data, they have been left out. Questions which are of relevance for Humanities scholars researching ancient languages have been added. The question have been answered to the best of the knowledge of the Daidalos… See the full description on the dataset page: https://huggingface.co/datasets/daidalos-project/latin_treebanks_ud_test.LatinEpicIntertextualityRetrieval
Latin Epic Intertextuality Retrieval
BEIR-style Retrieval benchmark for Latin epic intertextuality under PoetryMTEB.
Given a passage from Valerius Flaccus, Argonautica Book 1, retrieve the corresponding verse line(s) in Vergil (Aeneid), Lucan (Bellum Civile), Ovid (Metamorphoses), or Statius (Thebaid) that traditional scholarship identifies as parallels.
Gold parallels: Burns et al., NAACL-HLT 2021 (paper; repo), 945 curated pairs.
Dataset Card
Item… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/LatinEpicIntertextualityRetrieval.Tridis-latinLatinoamerican_Spanish_Voice_DatasetCompiled, and curated from the Crowdsourced high-quality speech datasets made by Google, available at https://openslr.org/resources.php.
This dataset consists of a wavs folder with the audios plus a .txt file with the path to the audio and the speaker's transcription.
Example:
wavs/vem_05223_00896110924.wav|Los corazones de pollo son una delicia.
wavs/vem_04310_01196944169.wav|Es un plato muy nutritivo.
wavs/vem_02484_00854567505.wav|En este momento estoy enviando a sus mails unos links de… See the full description on the dataset page: https://huggingface.co/datasets/Blakus/Latinoamerican_Spanish_Voice_Dataset.latin_italian_parallel
Italian-Latin Parallel Corpus (30,000 Sentences)
This dataset provides approximately 30,000 parallel sentences between Italian and Latin. It is designed for tasks such as machine translation and cross-linguistic research.
🌐 The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited.
Key Features
Size: ~30,000 translation pairs.
Languages: Italian (it), Latin (la).
Source… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin_italian_parallel.Kabyle-Latin-to-Tifinagh-Parallel-Corpus
Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.Latin-CC-170M
Latin-CC-170M
Reupload of the Corpus Corporum as Parquet, originally from
Kaggle,
and reuploaded as CSV by
Fece228/latin-literature-dataset-170M.
License
These works are public domain.
Original README
This is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin up to 19th century, which includes:
Classical Latin: works of Caesar, Cicero and many more
Medieval Latin: a… See the full description on the dataset page: https://huggingface.co/datasets/AncientLanguages/Latin-CC-170M.latin-vulgate
latin-vulgate
For machine learning translation of English-Latin and Latin-English — the complete Latin Vulgate Bible from the 4th Century AD and the complete English Douay-Rheims Bible from 1582.
Citation
If you use this data set, please support me by citing the repository. See APA Style.
APA Citation:
Flor, Nick. (2023). Latin Vulgate. GitHub. https://huggingface.co/datasets/professorf/latin-vulgate
MLA Citation:
Flor, Nick. Latin Vulgate. 2023, GitHub… See the full description on the dataset page: https://huggingface.co/datasets/professorf/latin-vulgate.kurdish-latin-wikipedia-sentences
Kurdish Latin Wikipedia Sentences Dataset
This dataset consists of 78,004 Kurdish sentences extracted from Wikipedia. All sentences are written in Latin script and consist of 12 to 18 words. The dataset has been carefully cleaned to remove numbers, dates, or non-textual elements.
Dataset Highlights
Source: Wikipedia (Kurdish content)
Script: Kurdish Latin
Content: Pure textual sentences (no numbers, dates, or special characters)
Sentence Length: 12 to 18 words per… See the full description on the dataset page: https://huggingface.co/datasets/zinaro/kurdish-latin-wikipedia-sentences.latin-greek-hebrew-english-dataset
Proverbs: Ancient Languages Set
This repository contains a collection of 2,000 short phrases translated into three ancient languages: Ancient Latin, Ancient Greek, Biblical Hebrew, and English. The phrases cover a wide variety of contexts, providing insight into the linguistic, cultural, and philosophical landscapes of these ancient civilizations.
Overview
The "Proverbs: Ancient Languages Set" is a resource designed to help individuals explore and understand ancient… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin-greek-hebrew-english-dataset.madlad-subset-latin-relatedcyrillic-vs-latin-tokenization
Cyrillic Tokenization Overhead Benchmark
This dataset accompanies the paper "Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems" submitted to the MRL Workshop at EMNLP 2026. It contains all the data needed to reproduce the paper's three studies, along with a balanced BPE tokenizer trained as part of the research.
What's inside
Directory
What it contains
study01_corpus_benchmark/
Tokenization fertility measured on the BrUK corpus (1.34M… See the full description on the dataset page: https://huggingface.co/datasets/idobrovolskyi/cyrillic-vs-latin-tokenization.latino40deita-10k-v0-sft-latinSame as HuggingFaceH4/deita-10k-v0-sft but without non-latin text.
