CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pstroe /cc100-latin Latin part of cc100 corpus This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface. Preprocessing I undertook the following preprocessing steps: Removal of all "pseudo-Latin" text ("Lorem ipsum ..."). Use of CLTK for sentence splitting and normalisation. Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.textn<1K9 likes155 downloads4y agoHugging Face02aimgo /Latin-OCR-Artifacts Latin sentences sourced from The Latin Library, converted to images, were subsequently degraded via OCRODEG. OCR (via Kraken and Tesseract) transcriptions were generated. The dataset was then augmented with several synthetic noise patterns in order to emulate the more severe corruption found in many older digitizations. If you use this in your work, please cite: @misc{mccarthy2025LACOROCR, author = {McCarthy, A. M.}, title = {{Latin OCR Artifacts}}, year = {2025}… See the full description on the dataset page: https://huggingface.co/datasets/aimgo/Latin-OCR-Artifacts.tabulartext-generation100K<n<1M0 likes36 downloads8mo agoHugging Face03medieval-data /medieval-latin-ner-HOME-AlcarThis NER dataset comes from the following publication. Stutzmann, D., Torres Aguilar, S., & Chaffenet, P. (2021). HOME-Alcar: Aligned and Annotated Cartularies [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5600884 I have used my the Notebook convert.ipynb to convert it from the original format to spaCy's format. The notebook expects the Database download from the link above to be in the root directory. The dataset contains nested spans. This is important because specific individual's… See the full description on the dataset page: https://huggingface.co/datasets/medieval-data/medieval-latin-ner-HOME-Alcar.text1K<n<10K1 likes30 downloads2y agoHugging Face04medieval-data /medieval-latin-ner-HOME-Alcar-sentsThis NER dataset comes from the following publication. Stutzmann, D., Torres Aguilar, S., & Chaffenet, P. (2021). HOME-Alcar: Aligned and Annotated Cartularies [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5600884 The dataset contains nested spans. This is important because specific individual's identities are often tied to a specific place. This means that the location entity is part of the person identity. To keep this nuance, the dataset should be used with the spaCy SpanCat pipe.… See the full description on the dataset page: https://huggingface.co/datasets/medieval-data/medieval-latin-ner-HOME-Alcar-sents.text10K<n<100K1 likes16 downloads2y agoHugging Face05lsb /ancient-latin-passagestextn<1K0 likes14 downloads5y agoHugging Face06irtizaab /Latin-Roman-Urdu-QnAtextn<1K0 likes5 downloads2y agoHugging Face07cudecanarim /8th-street-latinasimage10K<n<100K0 likes3 downloads1y agoHugging Face08elWaiEle /Las-Barras-Latinas-Mas-Cabronas-v1textn<1K0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.