CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ken-Z /Latin-Audio Dataset Summary Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training. Alignment and curation: Kaiyuan Zhao Language: Latin (Classical) Uses This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.audiotext-to-speech10K<n<100K8 likes3.2k downloads2mo agoHugging Face02PleIAs /Latin-PD 🇲🇪 Latin Public Domain Books (Latin) 🇲🇪 Latin-Public Domain or Latin-PD is a large collection aiming to aggregate all Latin monographies and periodicals in the public domain. As of June 2024, it is the largest Latin open corpus. Dataset summary The collection contains 16,521,454,086 words (159,070 titles) recovered from multiple sources, including the Internet Archive and various European national libraries and cultural heritage institutions (BDH, BNF). Each parquet… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Latin-PD.tabular100K<n<1M8 likes2.4k downloads2y agoHugging Face03Fece228 /latin-literature-dataset-170MThis is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin. The dataset is split in two parts: preprocessed with basic cltk tools, ready for work, and raw text data. It must be noted, however, that the latter contains text in Greek, Hebrew, and other languages, with references and contractions text100M<n<1B11 likes837 downloads4y agoHugging Face04TomasGuija /LatinFontsSVGs SVG Font Dataset Overview We present a curated dataset of vector font glyphs stored as SVG files, designed for research in generative modeling, structured vector graphics and typography synthesis. The dataset was created for the development and evaluation of our paper: DesigNet: Learning to Draw Vector Graphics as Designers Do Related Resources Paper (arXiv) : https://arxiv.org/abs/2604.06494 Code: https://github.com/TomasGuija/DesigNet… See the full description on the dataset page: https://huggingface.co/datasets/TomasGuija/LatinFontsSVGs.tabular1M<n<10M0 likes595 downloads5mo agoHugging Face05LatinNLP /LatinSummarizer LatinSummarizer Dataset Structure aligned_en_la_data_raw.csv aligned_en_la_data_cleaned.csv aligned_en_la_data_cleaned_with_stanza.csv concat_aligned_data.csv concat_cleaned.csv latin_wikipedia_cleaned.csv latin_wikipedia_raw.csv latin-literature-dataset-170M_raw_cleaned.csv latin-literature-dataset-170M_raw_cleaned_chunked.csv Elsa_aligned/ README.md Details aligned_en_la_data_raw.csv This dataset contains aligned Latin (la) - English (en)… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/LatinSummarizer.texttranslation1M<n<10M0 likes412 downloads2y agoHugging Face06KhalfounMehdi /arabic-latin-invoices-synthetic Synthetic Arabic/Latin Invoices — label-first multimodal dataset A large, perfectly-labeled synthetic invoice dataset for training multimodal invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin layouts, and all three numeral glyph systems. Every sample is generated label-first: a structured invoice record is sampled, then rendered to pixels via a headless browser, and bounding boxes are read back from the same DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.imageimage-to-text1K<n<10K1 likes227 downloads3mo agoHugging Face07TartarusXXX /uyghur-cv-latinaudio100K<n<1M1 likes207 downloads9mo agoHugging Face08grosenthal /latin_english_translation Dataset Card for "latin_english_parallel" 101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation. For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences. Each sample is annotated with the index and file (and therefore author/work) that the sample is from. If you find errors… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_translation.texttranslation100K<n<1M14 likes163 downloads3y agoHugging Face09pstroe /cc100-latin Latin part of cc100 corpus This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface. Preprocessing I undertook the following preprocessing steps: Removal of all "pseudo-Latin" text ("Lorem ipsum ..."). Use of CLTK for sentence splitting and normalisation. Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.textn<1K9 likes155 downloads4y agoHugging Face10julian-schelb /latin-classical-intertextuality-corpus Latin Classical Authors Corpus This dataset contains processed texts from classical Latin authors, serving as a retrieval corpus for intertextuality research. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Cicero and classical Latin literature. Related Datasets This corpus is part of the Latin Jerome Intertextuality collection: Queries: latin-classical-intertextuality-queries - The whole works of… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-corpus.texttext-retrieval10K<n<100K4 likes137 downloads1mo agoHugging Face11julian-schelb /latin-classical-intertextuality-labels Latin Jerome Intertextuality Labels This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.tabulartext-retrieval1K<n<10K2 likes121 downloads1mo agoHugging Face12julian-schelb /latin-classical-intertextuality-queries Latin Classical Intertextuality Queries This dataset contains query texts used for finding intertextual relationships with classical Latin authors. It comprises the whole works of Hieronymus (Jerome) and Lactantius, which are searched against a corpus of classical Latin literature. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Lactantius and classical Latin literature. Related Datasets This queries… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-queries.texttext-retrieval10K<n<100K2 likes113 downloads1mo agoHugging Face13wnkh /medieval-latinimage10K<n<100K0 likes105 downloads9mo agoHugging Face14grosenthal /latin_english_parallel Dataset Card for "latin_english_parallel" 101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation. For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences. Additionally, the English translations were both 1. copyrighted and 2. outdated. As such, we decided to modernize and… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_parallel.texttranslation100K<n<1M11 likes98 downloads3y agoHugging Face15daidalos-project /latin_treebanks_ud_testNOTE: This template for datasheets for ancient language data is based on the proposal by Gebru et al. 2021 https://arxiv.org/abs/1803.09010. The majority of questions is taken from there, but in a slightly rearanged order. However, as some questions are not relevant for historical data, they have been left out. Questions which are of relevance for Humanities scholars researching ancient languages have been added. The question have been answered to the best of the knowledge of the Daidalos… See the full description on the dataset page: https://huggingface.co/datasets/daidalos-project/latin_treebanks_ud_test.texttoken-classification1K<n<10K0 likes82 downloads1mo agoHugging Face16PoetryMTEB /LatinEpicIntertextualityRetrieval Latin Epic Intertextuality Retrieval BEIR-style Retrieval benchmark for Latin epic intertextuality under PoetryMTEB. Given a passage from Valerius Flaccus, Argonautica Book 1, retrieve the corresponding verse line(s) in Vergil (Aeneid), Lucan (Bellum Civile), Ovid (Metamorphoses), or Statius (Thebaid) that traditional scholarship identifies as parallels. Gold parallels: Burns et al., NAACL-HLT 2021 (paper; repo), 945 curated pairs. Dataset Card Item… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/LatinEpicIntertextualityRetrieval.tabulartext-retrieval10K<n<100K0 likes80 downloads1mo agoHugging Face17wandb /deita-10k-v0-sft-latinSame as HuggingFaceH4/deita-10k-v0-sft but without non-latin text. text10K<n<100K1 likes72 downloads3y agoHugging Face18wnkh /Tridis-latinimage100K<n<1M0 likes71 downloads9mo agoHugging Face19Dddixyy /latin_italian_parallel Italian-Latin Parallel Corpus (30,000 Sentences) This dataset provides approximately 30,000 parallel sentences between Italian and Latin. It is designed for tasks such as machine translation and cross-linguistic research. 🌐 The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited. Key Features Size: ~30,000 translation pairs. Languages: Italian (it), Latin (la). Source… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin_italian_parallel.texttranslation10K<n<100K3 likes67 downloads1y agoHugging Face20abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes57 downloads3mo agoHugging Face21AncientLanguages /Latin-CC-170M Latin-CC-170M Reupload of the Corpus Corporum as Parquet, originally from Kaggle, and reuploaded as CSV by Fece228/latin-literature-dataset-170M. License These works are public domain. Original README This is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin up to 19th century, which includes: Classical Latin: works of Caesar, Cicero and many more Medieval Latin: a… See the full description on the dataset page: https://huggingface.co/datasets/AncientLanguages/Latin-CC-170M.tabular1K<n<10K1 likes56 downloads9mo agoHugging Face22zinaro /kurdish-latin-wikipedia-sentences Kurdish Latin Wikipedia Sentences Dataset This dataset consists of 78,004 Kurdish sentences extracted from Wikipedia. All sentences are written in Latin script and consist of 12 to 18 words. The dataset has been carefully cleaned to remove numbers, dates, or non-textual elements. Dataset Highlights Source: Wikipedia (Kurdish content) Script: Kurdish Latin Content: Pure textual sentences (no numbers, dates, or special characters) Sentence Length: 12 to 18 words per… See the full description on the dataset page: https://huggingface.co/datasets/zinaro/kurdish-latin-wikipedia-sentences.text10K<n<100K0 likes48 downloads2y agoHugging Face23Dddixyy /latin-greek-hebrew-english-dataset Proverbs: Ancient Languages Set This repository contains a collection of 2,000 short phrases translated into three ancient languages: Ancient Latin, Ancient Greek, Biblical Hebrew, and English. The phrases cover a wide variety of contexts, providing insight into the linguistic, cultural, and philosophical landscapes of these ancient civilizations. Overview The "Proverbs: Ancient Languages Set" is a resource designed to help individuals explore and understand ancient… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin-greek-hebrew-english-dataset.texttranslation1K<n<10K1 likes44 downloads2y agoHugging Face24ccde /madlad-subset-latin-relatedtext10K<n<100K0 likes44 downloads2mo agoHugging Face25LatinNLP /LatinSummarizerDataset LatinSummarizer Dataset Overview The LatinSummarizerDataset is a structured dataset used in the GitHub Repository for Latin summarization and translation tasks. This dataset provides aligned English-Latin texts, extractive summaries, and pre-training prompts for fine-tuning models like mT5 for low-resource NLP applications. Structure The dataset is divided into two main phases: Pre-training Data: Includes aligned bilingual corpora, synthetic… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/LatinSummarizerDataset.texttranslation0 likes42 downloads2y agoHugging Face26DRDELATV /woman-latina-loraimagen<1K3 likes38 downloads1y agoHugging Face27aimgo /Latin-OCR-Artifacts Latin sentences sourced from The Latin Library, converted to images, were subsequently degraded via OCRODEG. OCR (via Kraken and Tesseract) transcriptions were generated. The dataset was then augmented with several synthetic noise patterns in order to emulate the more severe corruption found in many older digitizations. If you use this in your work, please cite: @misc{mccarthy2025LACOROCR, author = {McCarthy, A. M.}, title = {{Latin OCR Artifacts}}, year = {2025}… See the full description on the dataset page: https://huggingface.co/datasets/aimgo/Latin-OCR-Artifacts.tabulartext-generation100K<n<1M0 likes36 downloads8mo agoHugging Face28Dddixyy /latino_italiano_traduzioni_DIRETTEtexttranslation1K<n<10K1 likes35 downloads2y agoHugging Face29lunovian /vietnamese-nom-latin-translationtexttranslation1K<n<10K1 likes34 downloads2y agoHugging Face30abdelhaqueidali /Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset Dataset Card for Tachelhit Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Tachelhit language ($\text{Tacelḥit}$ / $\text{Tamazigt}$), pairing native Latin-based orthography with automated Amazigh script transliterations. It is built by processing clean source sentences through an algorithmic engine designed to handle phonetic mappings, manage contextual schwa distributions, and safeguard acronyms and foreign proper names. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset.texttranslation10K<n<100K0 likes34 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.