datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/scanned-images-dataset-for-ocr-and-vlm-finetuning.scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/prabhats0605/scanned-images-dataset-for-ocr-and-vlm-finetuning.analog_clocks_combinations_for_finetuning
Analog Clocks Combinations Dataset for Finetuning
This repository hosts a collection of 43,200 high-quality, synthetic images of analog clocks, generated for every possible hour, minute, and second in a 12-hour cycle, and for each of three clock types:
Base: normal clocks.
Distorted: dial with distorted shape.
Modified hands: hands with the same thickness and with an arrow.
The data is useful for training and benchmarking computer vision models on tasks like time recognition… See the full description on the dataset page: https://huggingface.co/datasets/migonsa/analog_clocks_combinations_for_finetuning.
