16th
Datasets
All datasets matching “16th”portuguese-handwriting-16th-19th
Portuguese Handwriting 16th-19th c.
This dataset consists of images and transcribed pages of Portuguese manuscripts from the 16th to the 19th centuries, primarily derived from the court records of the Portuguese Inquisition. It is designed for training and evaluating Handwritten Text Recognition (HTR) models.
Dataset Description
Overview
This is a collection of images and corresponding ground truth transcriptions (text) of historical Portuguese handwritten… See the full description on the dataset page: https://huggingface.co/datasets/SSamDav/portuguese-handwriting-16th-19th.Hanserezesse_2half_16th_50s_linesfromxml
Dataset Card for Hanserezesse_2half_16th_50s_linesfromxml
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 3,560 samples across 1 split(s).
Dataset Structure
Data Splits
train: 3,560 samples
Dataset Size
Approximate total size: 907.57 MB
Total samples: 3,560
Features
image: Image(mode=None, decode=False)
text: Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/ingala09/Hanserezesse_2half_16th_50s_linesfromxml.pythia-1.4B-tldr-sft_tldr_pythia_1.4b_rm_sft_tldr_pythia_1.4b_16th_3_iter_1Hanserezesse_2half_16th_50s_rawxml
Dataset Card for Hanserezesse_2half_16th_50s_rawxml
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 1,000 samples across 1 split(s).
Dataset Structure
Data Splits
train: 1,000 samples
Dataset Size
Approximate total size: 8,737.32 MB
Total samples: 1,000
Features
image: Image(mode=None, decode=False)
xml_content: Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/ingala09/Hanserezesse_2half_16th_50s_rawxml.Hanserezesse_2half_16th_50s_lines
Dataset Card for Hanserezesse_2half_16th_50s_lines
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 3,560 samples across 1 split(s).
Dataset Structure
Data Splits
train: 3,560 samples
Dataset Size
Approximate total size: 554.70 MB
Total samples: 3,560
Features
image: Image(mode=None, decode=False)
text: Value('string')
line_id:… See the full description on the dataset page: https://huggingface.co/datasets/ingala09/Hanserezesse_2half_16th_50s_lines.pythia-1.4B-tldr-sft_tldr_pythia_1.4b_rm_sft_tldr_pythia_1.4b_16th_2_iter_1
