CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ken-Z /Latin-Audio Dataset Summary Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training. Alignment and curation: Kaiyuan Zhao Language: Latin (Classical) Uses This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.audiotext-to-speech10K<n<100K8 likes3.1k downloads2mo agoHugging Face02PleIAs /Latin-PD 🇲🇪 Latin Public Domain Books (Latin) 🇲🇪 Latin-Public Domain or Latin-PD is a large collection aiming to aggregate all Latin monographies and periodicals in the public domain. As of June 2024, it is the largest Latin open corpus. Dataset summary The collection contains 16,521,454,086 words (159,070 titles) recovered from multiple sources, including the Internet Archive and various European national libraries and cultural heritage institutions (BDH, BNF). Each parquet… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Latin-PD.tabular100K<n<1M8 likes2.4k downloads2y agoHugging Face03TomasGuija /LatinFontsSVGs SVG Font Dataset Overview We present a curated dataset of vector font glyphs stored as SVG files, designed for research in generative modeling, structured vector graphics and typography synthesis. The dataset was created for the development and evaluation of our paper: DesigNet: Learning to Draw Vector Graphics as Designers Do Related Resources Paper (arXiv) : https://arxiv.org/abs/2604.06494 Code: https://github.com/TomasGuija/DesigNet… See the full description on the dataset page: https://huggingface.co/datasets/TomasGuija/LatinFontsSVGs.tabular1M<n<10M0 likes1.3k downloads5mo agoHugging Face04Fece228 /latin-literature-dataset-170MThis is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin. The dataset is split in two parts: preprocessed with basic cltk tools, ready for work, and raw text data. It must be noted, however, that the latter contains text in Greek, Hebrew, and other languages, with references and contractions text100M<n<1B11 likes779 downloads4y agoHugging Face05LatinNLP /latin-summarizer-dataset ✨ LatinSummarizer Dataset ✨ Note: If Dataset Viewer is not available, see samples of dataset for samples from the dataset.The LatinSummarizer Dataset is a comprehensive collection of Latin texts designed to support natural language processing research for a low-resource language. It provides parallel data for various tasks, including translation (Latin-to-English) and summarization (extractive and abstractive). This dataset was created for a… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/latin-summarizer-dataset.summarization0 likes722 downloads1y agoHugging Face06LatinNLP /LatinSummarizer LatinSummarizer Dataset Structure aligned_en_la_data_raw.csv aligned_en_la_data_cleaned.csv aligned_en_la_data_cleaned_with_stanza.csv concat_aligned_data.csv concat_cleaned.csv latin_wikipedia_cleaned.csv latin_wikipedia_raw.csv latin-literature-dataset-170M_raw_cleaned.csv latin-literature-dataset-170M_raw_cleaned_chunked.csv Elsa_aligned/ README.md Details aligned_en_la_data_raw.csv This dataset contains aligned Latin (la) - English (en)… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/LatinSummarizer.texttranslation1M<n<10M0 likes408 downloads2y agoHugging Face07KhalfounMehdi /arabic-latin-invoices-synthetic Synthetic Arabic/Latin Invoices — label-first multimodal dataset A large, perfectly-labeled synthetic invoice dataset for training multimodal invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin layouts, and all three numeral glyph systems. Every sample is generated label-first: a structured invoice record is sampled, then rendered to pixels via a headless browser, and bounding boxes are read back from the same DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.imageimage-to-text1K<n<10K1 likes222 downloads3mo agoHugging Face08TartarusXXX /uyghur-cv-latinaudio100K<n<1M1 likes210 downloads9mo agoHugging Face09grosenthal /latin_english_translation Dataset Card for "latin_english_parallel" 101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation. For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences. Each sample is annotated with the index and file (and therefore author/work) that the sample is from. If you find errors… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_translation.texttranslation100K<n<1M14 likes162 downloads3y agoHugging Face10pstroe /cc100-latin Latin part of cc100 corpus This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface. Preprocessing I undertook the following preprocessing steps: Removal of all "pseudo-Latin" text ("Lorem ipsum ..."). Use of CLTK for sentence splitting and normalisation. Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.textn<1K9 likes155 downloads4y agoHugging Face11julian-schelb /latin-classical-intertextuality-corpus Latin Classical Authors Corpus This dataset contains processed texts from classical Latin authors, serving as a retrieval corpus for intertextuality research. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Cicero and classical Latin literature. Related Datasets This corpus is part of the Latin Jerome Intertextuality collection: Queries: latin-classical-intertextuality-queries - The whole works of… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-corpus.texttext-retrieval10K<n<100K4 likes152 downloads29d agoHugging Face12julian-schelb /latin-classical-intertextuality-labels Latin Jerome Intertextuality Labels This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.tabulartext-retrieval1K<n<10K2 likes136 downloads29d agoHugging Face13julian-schelb /latin-classical-intertextuality-queries Latin Classical Intertextuality Queries This dataset contains query texts used for finding intertextual relationships with classical Latin authors. It comprises the whole works of Hieronymus (Jerome) and Lactantius, which are searched against a corpus of classical Latin literature. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Lactantius and classical Latin literature. Related Datasets This queries… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-queries.texttext-retrieval10K<n<100K2 likes128 downloads29d agoHugging Face14itserr /WP8-Latin-Embeddings-Indices0 likes124 downloads1y agoHugging Face15wnkh /medieval-latinimage10K<n<100K0 likes106 downloads9mo agoHugging Face16grosenthal /latin_english_parallel Dataset Card for "latin_english_parallel" 101k translation pairs between Latin and English, split 99/1/1 as train/test/val. These have been collected roughly 66% from the Loeb Classical Library and 34% from the Vulgate translation. For those that were gathered from the Loeb Classical Library, alignment was performd manually between Source and Target sequences. Additionally, the English translations were both 1. copyrighted and 2. outdated. As such, we decided to modernize and… See the full description on the dataset page: https://huggingface.co/datasets/grosenthal/latin_english_parallel.texttranslation100K<n<1M11 likes96 downloads3y agoHugging Face17daidalos-project /latin_treebanks_ud_testNOTE: This template for datasheets for ancient language data is based on the proposal by Gebru et al. 2021 https://arxiv.org/abs/1803.09010. The majority of questions is taken from there, but in a slightly rearanged order. However, as some questions are not relevant for historical data, they have been left out. Questions which are of relevance for Humanities scholars researching ancient languages have been added. The question have been answered to the best of the knowledge of the Daidalos… See the full description on the dataset page: https://huggingface.co/datasets/daidalos-project/latin_treebanks_ud_test.texttoken-classification1K<n<10K0 likes81 downloads1mo agoHugging Face18PoetryMTEB /LatinEpicIntertextualityRetrieval Latin Epic Intertextuality Retrieval BEIR-style Retrieval benchmark for Latin epic intertextuality under PoetryMTEB. Given a passage from Valerius Flaccus, Argonautica Book 1, retrieve the corresponding verse line(s) in Vergil (Aeneid), Lucan (Bellum Civile), Ovid (Metamorphoses), or Statius (Thebaid) that traditional scholarship identifies as parallels. Gold parallels: Burns et al., NAACL-HLT 2021 (paper; repo), 945 curated pairs. Dataset Card Item… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/LatinEpicIntertextualityRetrieval.tabulartext-retrieval10K<n<100K0 likes79 downloads1mo agoHugging Face19wnkh /Tridis-latinimage100K<n<1M0 likes75 downloads9mo agoHugging Face20Blakus /Latinoamerican_Spanish_Voice_DatasetCompiled, and curated from the Crowdsourced high-quality speech datasets made by Google, available at https://openslr.org/resources.php. This dataset consists of a wavs folder with the audios plus a .txt file with the path to the audio and the speaker's transcription. Example: wavs/vem_05223_00896110924.wav|Los corazones de pollo son una delicia. wavs/vem_04310_01196944169.wav|Es un plato muy nutritivo. wavs/vem_02484_00854567505.wav|En este momento estoy enviando a sus mails unos links de… See the full description on the dataset page: https://huggingface.co/datasets/Blakus/Latinoamerican_Spanish_Voice_Dataset.text-to-speech1 likes64 downloads2y agoHugging Face21Dddixyy /latin_italian_parallel Italian-Latin Parallel Corpus (30,000 Sentences) This dataset provides approximately 30,000 parallel sentences between Italian and Latin. It is designed for tasks such as machine translation and cross-linguistic research. 🌐 The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited. Key Features Size: ~30,000 translation pairs. Languages: Italian (it), Latin (la). Source… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin_italian_parallel.texttranslation10K<n<100K3 likes61 downloads1y agoHugging Face22abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes57 downloads3mo agoHugging Face23AncientLanguages /Latin-CC-170M Latin-CC-170M Reupload of the Corpus Corporum as Parquet, originally from Kaggle, and reuploaded as CSV by Fece228/latin-literature-dataset-170M. License These works are public domain. Original README This is a dataset collected from all the texts available at Corpus Corporum, which includes probably all the literary works ever written in Latin up to 19th century, which includes: Classical Latin: works of Caesar, Cicero and many more Medieval Latin: a… See the full description on the dataset page: https://huggingface.co/datasets/AncientLanguages/Latin-CC-170M.tabular1K<n<10K1 likes56 downloads9mo agoHugging Face24professorf /latin-vulgate latin-vulgate For machine learning translation of English-Latin and Latin-English — the complete Latin Vulgate Bible from the 4th Century AD and the complete English Douay-Rheims Bible from 1582. Citation If you use this data set, please support me by citing the repository. See APA Style. APA Citation: Flor, Nick. (2023). Latin Vulgate. GitHub. https://huggingface.co/datasets/professorf/latin-vulgate MLA Citation: Flor, Nick. Latin Vulgate. 2023, GitHub… See the full description on the dataset page: https://huggingface.co/datasets/professorf/latin-vulgate.2 likes53 downloads2y agoHugging Face25zinaro /kurdish-latin-wikipedia-sentences Kurdish Latin Wikipedia Sentences Dataset This dataset consists of 78,004 Kurdish sentences extracted from Wikipedia. All sentences are written in Latin script and consist of 12 to 18 words. The dataset has been carefully cleaned to remove numbers, dates, or non-textual elements. Dataset Highlights Source: Wikipedia (Kurdish content) Script: Kurdish Latin Content: Pure textual sentences (no numbers, dates, or special characters) Sentence Length: 12 to 18 words per… See the full description on the dataset page: https://huggingface.co/datasets/zinaro/kurdish-latin-wikipedia-sentences.text10K<n<100K0 likes49 downloads2y agoHugging Face26Dddixyy /latin-greek-hebrew-english-dataset Proverbs: Ancient Languages Set This repository contains a collection of 2,000 short phrases translated into three ancient languages: Ancient Latin, Ancient Greek, Biblical Hebrew, and English. The phrases cover a wide variety of contexts, providing insight into the linguistic, cultural, and philosophical landscapes of these ancient civilizations. Overview The "Proverbs: Ancient Languages Set" is a resource designed to help individuals explore and understand ancient… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/latin-greek-hebrew-english-dataset.texttranslation1K<n<10K1 likes44 downloads2y agoHugging Face27ccde /madlad-subset-latin-relatedtext10K<n<100K0 likes44 downloads2mo agoHugging Face28idobrovolskyi /cyrillic-vs-latin-tokenization Cyrillic Tokenization Overhead Benchmark This dataset accompanies the paper "Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems" submitted to the MRL Workshop at EMNLP 2026. It contains all the data needed to reproduce the paper's three studies, along with a balanced BPE tokenizer trained as part of the research. What's inside Directory What it contains study01_corpus_benchmark/ Tokenization fertility measured on the BrUK corpus (1.34M… See the full description on the dataset page: https://huggingface.co/datasets/idobrovolskyi/cyrillic-vs-latin-tokenization.text-classification1K<n<10K0 likes42 downloads3mo agoHugging Face29cristinakuo /latino400 likes40 downloads5y agoHugging Face30wandb /deita-10k-v0-sft-latinSame as HuggingFaceH4/deita-10k-v0-sft but without non-latin text. text10K<n<100K1 likes40 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.