datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese-ocr-datasettask_categories:
image-to-text
task_ids:
optical-character-recognition
text-recognition
Portuguese OCR Dataset
A comprehensive dataset for Portuguese OCR (Optical Character Recognition) generated from classic Portuguese literature with diverse fonts and visual styles.
Dataset Description
This dataset contains 20000 text images for OCR training, created from Portuguese books from Project Gutenberg. Each image contains a complete Portuguese sentence with proper… See the full description on the dataset page: https://huggingface.co/datasets/mazafard/portuguese-ocr-dataset.gpt4v-dataset-portugueseportuguese-ocr-dataset
