CoolFace
Datasetpublic

caveman273/aida-typewritten

typewritten OCR training data from AIDA-project Dataset Summary This dataset contains typewritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality typwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported Tasks The dataset was created for… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-typewritten.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes28downloads
5 commits on main
136485f5mo ago

Upload README.md with huggingface_hub

caveman273
13a40b85mo ago

Upload validation.parquet with huggingface_hub

caveman273
98b2dc35mo ago

Upload test.parquet with huggingface_hub

caveman273
4832d205mo ago

Upload train.parquet with huggingface_hub

caveman273
7c95e1f5mo ago

initial commit

caveman273