CoolFace
20 results

caveman

caveman273 /aida-handwritten Handwritten OCR training data from AIDA-project Dataset Summary This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported Tasks The dataset was created for… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-handwritten.imageimage-to-text1K<n<10K1 likes84 downloads5mo agoHugging Facecatsaresupercool /synthetic-caveman-thinkingtext1K<n<10K3 likes40 downloads3mo agoHugging Facecaveman273 /aida-ship-info Handwritten OCR training data from AIDA-project (Ship Registry) Dataset Summary This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations from ship registry records — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-ship-info.imageimage-to-text1K<n<10K0 likes37 downloads5mo agoHugging Facecaveman273 /aida-typewritten typewritten OCR training data from AIDA-project Dataset Summary This dataset contains typewritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality typwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported Tasks The dataset was created for optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-typewritten.imageimage-to-text10K<n<100K0 likes28 downloads5mo agoHugging Facenibauman /objectnav-sft-claude-cavemanimagen<1K1 likes28 downloads4mo agoHugging Facecaveman273 /theseus_ocr_tiny Theseus Finnish OCR Dataset Paragraph-level OCR dataset harvested from Theseus.fi, the Finnish repository of university of applied sciences theses. Each record is one paragraph crop extracted from a thesis PDF, paired with the text extracted by pdfplumber. Image Resolution Paragraph crops are rendered at 300 DPI (dots per inch) with 2 px padding on each side. At 300 DPI a standard A4 page is 2481 × 3507 pixels, giving high enough resolution for training OCR and… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/theseus_ocr_tiny.imageimage-to-text10K<n<100K0 likes22 downloads5mo agoHugging Face