lines
Datasets
All datasets matching “lines”glyph_machina_medieval_lines
glyph_machina_medieval_lines
Noisy HTR pretraining set: text-line crops from pre-Elizabeth-I English legal
manuscripts (AALT scans), with machine-generated transcriptions (confidence
prefixes stripped, confidence-filtered upstream). Line images are dewarped,
background-subtracted, inverted, 64 px tall.
Format: page-grouped WebDataset
data/*.tar are WebDataset shards (~1 GB each). One sample = one page.
For a page whose key is e.g.… See the full description on the dataset page: https://huggingface.co/datasets/mzzhang2014/glyph_machina_medieval_lines.c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines
Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines"
More Information needed
ocr-pl-lines
ocr-pl-lines
Syntetyczny zbiór linii tekstu po polsku do fine-tuningu OCR (TrOCR).
Pary NNNNN.png (obraz linii) + NNNNN.txt (transkrypcja).
Struktura
train/ — 2000 par (seed 42)
val/ — 200 par (seed 123)
Generowanie
OCR_engine —
python -m training.generate_synthetic
Korpus: zdania potoczne i urzędowe, domeny (faktury, umowy, medyczne,
prawnicze), losowe daty/kwoty/adresy/NIP/PESEL, zdania z pl.wikipedia.org.
Augmentacje: pochylenie, blur, szum… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ocr-pl-lines.simpsons_script_linesjam-alt-linessimpsons_script_lines_parsed
