caveman
Datasets
All datasets matching “caveman”aida-handwritten
Handwritten OCR training data from AIDA-project
Dataset Summary
This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported Tasks
The dataset was created for… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-handwritten.synthetic-caveman-thinkingaida-ship-info
Handwritten OCR training data from AIDA-project (Ship Registry)
Dataset Summary
This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations from ship registry records — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-ship-info.aida-typewritten
typewritten OCR training data from AIDA-project
Dataset Summary
This dataset contains typewritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality typwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported Tasks
The dataset was created for optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-typewritten.objectnav-sft-claude-cavemantheseus_ocr_tiny
Theseus Finnish OCR Dataset
Paragraph-level OCR dataset harvested from Theseus.fi,
the Finnish repository of university of applied sciences theses.
Each record is one paragraph crop extracted from a thesis PDF, paired with the
text extracted by pdfplumber.
Image Resolution
Paragraph crops are rendered at 300 DPI (dots per inch) with 2 px
padding on each side. At 300 DPI a standard A4 page is
2481 × 3507 pixels, giving high enough resolution for training
OCR and… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/theseus_ocr_tiny.
