datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc100-latin
Latin part of cc100 corpus
This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface.
Preprocessing
I undertook the following preprocessing steps:
Removal of all "pseudo-Latin" text ("Lorem ipsum ...").
Use of CLTK for sentence splitting and normalisation.
Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.Latin-OCR-Artifacts
Latin sentences sourced from The Latin Library, converted to images, were subsequently degraded via OCRODEG.
OCR (via Kraken and Tesseract) transcriptions were generated. The dataset was then augmented with several synthetic noise patterns in order to emulate the more severe corruption found in many older digitizations.
If you use this in your work, please cite:
@misc{mccarthy2025LACOROCR,
author = {McCarthy, A. M.},
title = {{Latin OCR Artifacts}},
year = {2025}… See the full description on the dataset page: https://huggingface.co/datasets/aimgo/Latin-OCR-Artifacts.medieval-latin-ner-HOME-AlcarThis NER dataset comes from the following publication.
Stutzmann, D., Torres Aguilar, S., & Chaffenet, P. (2021). HOME-Alcar: Aligned and Annotated Cartularies [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5600884
I have used my the Notebook convert.ipynb to convert it from the original format to spaCy's format. The notebook expects the Database download from the link above to be in the root directory.
The dataset contains nested spans. This is important because specific individual's… See the full description on the dataset page: https://huggingface.co/datasets/medieval-data/medieval-latin-ner-HOME-Alcar.medieval-latin-ner-HOME-Alcar-sentsThis NER dataset comes from the following publication.
Stutzmann, D., Torres Aguilar, S., & Chaffenet, P. (2021). HOME-Alcar: Aligned and Annotated Cartularies [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5600884
The dataset contains nested spans. This is important because specific individual's identities are often tied to a specific place. This means that the location entity is part of the person identity. To keep this nuance, the dataset should be used with the spaCy SpanCat pipe.… See the full description on the dataset page: https://huggingface.co/datasets/medieval-data/medieval-latin-ner-HOME-Alcar-sents.ancient-latin-passagesLatin-Roman-Urdu-QnA8th-street-latinasLas-Barras-Latinas-Mas-Cabronas-v1
