datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
handwritten_text_detection
Handwritten text detection dataset
Data domain
The blanks were provided by youth organization "Armenian Club" (telegram, instagram ), Russia Moscow.
The text on blanks was written during dictation "Teladrutyun" in 2018
The blanks were labeled by Amir and Renal during research project in HSE MIEM
Dataset info
Contains labeled dictations blanks in YOLO format
91 image in total, 73 (80%) for train and 18 (20%) for test
No image alignment or any preprocess… See the full description on the dataset page: https://huggingface.co/datasets/armvectores/handwritten_text_detection.Nepali_Text_Detection_Datasetocr-text-detection-in-the-documents
OCR Text Detection in the Documents Object Detection dataset
The dataset is a collection of images that have been annotated with the location of text in the document. The dataset is specifically curated for text detection and recognition tasks in documents such as scanned papers, forms, invoices, and handwritten notes.
The dataset contains a variety of document types, including different layouts, font sizes, and styles. The images come from diverse sources, ensuring a representative… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/ocr-text-detection-in-the-documents.receipts-text-detectionmanga-covers-text-detection
Manga Covers Text Detection
Manually annotated text boxes for manga covers text detection created by @JustANormalTinkerer.
This dataset contains 100 cover images and 526 manually annotated text rectangles. The image column is a Hugging Face image feature containing the original image bytes. The boxes column contains normalized pixel-coordinate bounding boxes in [x_min, y_min, x_max, y_max] format, the original points, label, and shape metadata. annotation_json preserves the… See the full description on the dataset page: https://huggingface.co/datasets/petersunde/manga-covers-text-detection.ocr-generated-machine-readable-zone-mrz-text-detection
OCR GENERATED Machine-Readable Zone (MRZ) Text Detection
The dataset includes a collection of GENERATED photos containing Machine Readable Zones (MRZ) commonly found on identification documents such as passports, visas, and ID cards. Each photo in the dataset is accompanied by text detection and Optical Character Recognition (OCR) results.
💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the dataset, leave a request on our website to… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/ocr-generated-machine-readable-zone-mrz-text-detection.EvArEST-dataset-for-Arabic-scene-text-detection
EvArEST
Everyday Arabic-English Scene Text dataset, from the paper: Arabic Scene Text Recognition in the Deep Learning Era: Analysis on A Novel Dataset
Detection Dataset
The text detection dataset has 510 images all containing one or more instances of text. Each word is annotated with a four-point polygon that starts with the top left corner of the polygon and follows clockwise. Each image comes with a text file containing three attributes: the four points of the polygon… See the full description on the dataset page: https://huggingface.co/datasets/Melaraby/EvArEST-dataset-for-Arabic-scene-text-detection.handwritten_text_detection
Handwritten text detection dataset
Data domain
The blanks were provided by youth organization "Armenian Club" (telegram, instagram ), Russia Moscow.
The text on blanks was written during dictation "Teladrutyun" in 2018
The blanks were labeled by Amir and Renal during research project in HSE MIEM
Dataset info
Contains labeled dictations blanks in YOLO format
91 image in total, 73 (80%) for train and 18 (20%) for test
No image alignment or any… See the full description on the dataset page: https://huggingface.co/datasets/akgupta4332/handwritten_text_detection.Khmer-Text-Detection-0.2k
SoyVitou/Khmer-Text-Detection-0.2k
Khmer OCR dataset for scene text detection + transcription.
This dataset is packaged as a Hugging Face dataset using a single Parquet file:
train/metadata.parquet
✅ The image column is stored as embedded bytes inside the Parquet, so load_dataset() works without downloading a separate images folder.
Dataset format
Each row contains:
id (string): sample id
image (image): image object (decoded by datasets)
annotation (string):… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/Khmer-Text-Detection-0.2k.experimento3-industrial-text-detection
Experimento-3 - Industrial Machinery Text Detection Dataset
Dataset Description
This dataset contains 4,237 images of industrial machinery nameplates with detailed text field annotations for OCR and information extraction tasks. The dataset focuses on extracting key information from equipment nameplates including manufacturer, model, serial numbers, and dates.
Dataset Summary
Task: Industrial text detection and OCR
Domain: Industrial machinery and equipment… See the full description on the dataset page: https://huggingface.co/datasets/kahua-ml/experimento3-industrial-text-detection.TextDetectionData
Text Detection Dataset
This dataset is a comprehensive collection of Arabic and English text detection samples designed for benchmarking and evaluating text detection models. The dataset combines samples from multiple open-source datasets to provide diverse text detection challenges across different domains, scripts, and image conditions.
📁 Dataset Structure
Images are stored in .jpg format in the "Text Detection Images" folder
Accompanied by a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/BatSilver/TextDetectionData.text_detections_easyocr_yolov10
