datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
latin-asr-post-processing-dataset
Latin ASR Post-Processing Dataset
A sequence-labeling dataset built for fine-tuning BERT-style models (e.g., latin-bert) on Inverse Text Normalization (ITN) — restoring capitalization and punctuation on raw, lowercased Latin text (such as njand/wav2vec2-xls-r-latin ASR outputs).
Compiled from 2,141 files in the CLTK Latin Library and augmented with transcripts from the njand/llpsi-speech-dataset (currently private). Cleaned and transformed through a specialized classical Latin… See the full description on the dataset page: https://huggingface.co/datasets/njand/latin-asr-post-processing-dataset.ASR_Post-processing-evalASR_Post-processing-dataset
