datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LatinFontsSVGs
SVG Font Dataset
Overview
We present a curated dataset of vector font glyphs stored as SVG files, designed for research in generative modeling, structured vector graphics and typography synthesis.
The dataset was created for the development and evaluation of our paper:
DesigNet: Learning to Draw Vector Graphics as Designers Do
Related Resources
Paper (arXiv) : https://arxiv.org/abs/2604.06494
Code: https://github.com/TomasGuija/DesigNet… See the full description on the dataset page: https://huggingface.co/datasets/TomasGuija/LatinFontsSVGs.LatinSummarizer
LatinSummarizer Dataset
Structure
aligned_en_la_data_raw.csv
aligned_en_la_data_cleaned.csv
aligned_en_la_data_cleaned_with_stanza.csv
concat_aligned_data.csv
concat_cleaned.csv
latin_wikipedia_cleaned.csv
latin_wikipedia_raw.csv
latin-literature-dataset-170M_raw_cleaned.csv
latin-literature-dataset-170M_raw_cleaned_chunked.csv
Elsa_aligned/
README.md
Details
aligned_en_la_data_raw.csv
This dataset contains aligned Latin (la) - English (en)… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/LatinSummarizer.Kabyle-Latin-to-Tifinagh-Parallel-Corpus
Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.vietnamese-nom-latin-translationTashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset
Dataset Card for Tachelhit Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Tachelhit language ($\text{Tacelḥit}$ / $\text{Tamazigt}$), pairing native Latin-based orthography with automated Amazigh script transliterations. It is built by processing clean source sentences through an algorithmic engine designed to handle phonetic mappings, manage contextual schwa distributions, and safeguard acronyms and foreign proper names.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset.latin-intertextuality-modification-test-cases
Latin Intertextuality Test Cases (Controlled Modifications)
This repository provides a small qualitative dataset of Latin sentence pairs used to study how sentence-embedding similarity reacts to controlled modifications of a query sentence in intertextual links (literal quotation, paraphrase, allusion).
The dataset contains 11 base cases, each paired with 12 modified variants of the query sentence (plus one random-control pairing), for a total of 143 rows.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/MichaelWittweiler/latin-intertextuality-modification-test-cases.NLLB-Seed_Tamasheq-Latin-Scripten-sat-latin-splitgreek_latin_authors
Greek and Latin Authors
This dataset contains the names of authors who primarily wrote in Ancient Greek (labeled "Greek")
and authors who primarily wrote in Latin (labeled "Latin").
The Greek names were gathered from the Thesaurus Linguae Graecae (TLG) project.
Specifically, the TLG makes lists of authors openly available at https://stephanus.tlg.uci.edu/tlgauthors/post_tlg_e.php
and https://stephanus.tlg.uci.edu/tlgauthors/cd.authors.php. These names were supplemented by names… See the full description on the dataset page: https://huggingface.co/datasets/sjhuskey/greek_latin_authors.latin_author_dll_id
Latin Author Name to Digital Latin Library ID Concordance
This dataset contains variant names of authors of works in Latin and their corresponding identifier in the Digital Latin Library's Catalog.
The variant names were gathered from the Virtual International Authority File records for the authors. In instances where an author's name has few or no known variant name forms, pseudo-variant names were generated by "misspelling" the name using the following script:
from textblob import… See the full description on the dataset page: https://huggingface.co/datasets/sjhuskey/latin_author_dll_id.latin_italian_semantic_translation
Latin - Italian (Semantic Context And Translation)
This dataset was generated by fetching random first paragraphs from Wikipedia in Ancient Latin (la.wikipedia.org)
and processing them using Gemini AI with the following goal(s):
Processing Goal 1: mantiene il testo in latino. ma riduce l'ambiguità aggiungendo tag grammaticali IN ITALIANO (come soggetto, verbo eccetera...) e tag funzionali (come "indica dove è nato il soggetto" oppure "indica che il soggetto possiede l'oggetto")… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin_italian_semantic_translation.english-darija-latin-mergedItalian_latin_parallel_animals
descrizioni di animali e habitat - Synthetic Dataset
This dataset was generated using the Synthetic Dataset Generator powered by Gemini AI.
Topic: descrizioni di animali e habitat
Field 1: italiano
Field 2: latino antico(traduzione)
Rows: 280
Generated on: 2025-05-27T00:07:49.042Z
latin_lyrics
