datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
handwritten-dates-numbers-ocrreddit-olympics-2024DiEm_HTR-Numbers
Dataset Card for DiEm HTR Numbers
The DiEm HTR Numbers dataset is a ground truth dataset consisting of numbers written in historical danish handwriting from the 18th century, generated as part of the Digitalisering af Enesteministerialbøger project at the Danish National Archives.
Dataset Details
Dataset Description
The Digitalisering af Enesteministerialbøger project (DiEm) at the Danish National Archives aims to transcribe and make publically available… See the full description on the dataset page: https://huggingface.co/datasets/RA-Data-Science/DiEm_HTR-Numbers.amount-in-numbers-dzNumbers-Sign-Language-Datasetmnist-numbers-0to10000-128x128
MNIST Numbers 0..10,000 (128×128)
10,000 synthetic grayscale images composed from MNIST digits (black on white), resized to 128×128.
Each row corresponds to an integer n ∈ [0, 10,000] and includes:
image: digits tiled left→right with small rotation jitter
digits: e.g., "10000"
words: e.g., "ten thousand" (no "and")
value: integer 0..10,000
length: number of digits (1..5)
Splits
train: 9,000
test: 1,000
Usage
from datasets import load_dataset
DS =… See the full description on the dataset page: https://huggingface.co/datasets/starkdv123/mnist-numbers-0to10000-128x128.OCR-Numbers-Printed-0A synthetic dataset for text recognition tasks, contains 300,000 images, numbers only.
OCR-Numbers-Printed-5A synthetic dataset for text recognition tasks, contains 100,000 images, numbers only.
numbers-datasetwiki_numbers_ocrOCR-Numbers-Printed-2A synthetic dataset for text recognition tasks, contains 300,000 images.
vqa_v2_numbers_promptedvqa_v2_numbers_prompted-v2
