handwritten-text
handwritten_text_detection
Handwritten text detection dataset
Data domain
The blanks were provided by youth organization "Armenian Club" (telegram, instagram ), Russia Moscow.
The text on blanks was written during dictation "Teladrutyun" in 2018
The blanks were labeled by Amir and Renal during research project in HSE MIEM
Dataset info
Contains labeled dictations blanks in YOLO format
91 image in total, 73 (80%) for train and 18 (20%) for test
No image alignment or any preprocess… See the full description on the dataset page: https://huggingface.co/datasets/armvectores/handwritten_text_detection.image-text_handwritten-bundesratsprotokolle_xix-xx!!!This data set does not contain Ground Truth!!!
--- Data has been automatically created, using ATR models ---
Dataset Card for transkribus-exports-74823-raw-xml
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 148494 samples across 1 split(s).
These are ''automatically'' transcribed pages.
Images provided by the Federal Archives (Schweizerisches Bundesarchiv, BAR)
For a description of the project… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_handwritten-bundesratsprotokolle_xix-xx.FrenchCensus-handwritten-texts
Source
This repository contains 3 datasets created within the POPP project (Project for the Oceration of the Paris Population Census) for the task of handwriting text recognition. These datasets have been published in Recognition and information extraction in historical handwritten tables: toward understanding early 20th century Paris census at DAS 2022.
The 3 datasets are called “Generic dataset”, “Belleville”, and “Chaussée d’Antin” and contains lines made from the extracted rows… See the full description on the dataset page: https://huggingface.co/datasets/agomberto/FrenchCensus-handwritten-texts.handwritten-texthandwritten_text_ocrArabic-Handwritten-Text-Recognition-Dataset
Dataset Description
This dataset is a re-uploaded version of the Muharaf dataset.
The original dataset was created by Mehreen Saeed et al. and released for research purposes.
This repository is intended for easier access and experimentation via Hugging Face.
|
How to use
from datasets import load_dataset
ds = load_dataset("amjad-awad/Arabic-Handwritten-Text-Recognition-Dataset")
print(ds["train"][0]["image"])
Attribution
All credit for creating and… See the full description on the dataset page: https://huggingface.co/datasets/amjad-awad/Arabic-Handwritten-Text-Recognition-Dataset.
