CoolFace
Datasetpublic

techiaith/finepdfs-cy-errors

Dataset Card: finepdfs-cy-errors Description This dataset contains Welsh-language text extracted from PDFs using rolmOCR, with automated spelling and grammar error annotations generated by Cysill (the Welsh spell checker). The dataset is derived from the Welsh (cym_Latn) subset of HuggingFaceFW/finepdfs, filtered to include only documents processed with the rolmOCR extractor. Dataset Statistics Corpus Size Total number of words: 1… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/finepdfs-cy-errors.

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes13downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
techiaith/finepdfs-cy-errors · CoolFace