Teklia/NewsEye-Austrian-line
NewsEye Austrian - line level Dataset Summary The dataset comprises Austrian newspaper pages from 19th and early 20th century. The images were provided by the Austrian National Library. Languages The documents are in Austrian German with the Fraktur font. Note that all images are resized to a fixed height of 128 pixels. Dataset Structure Data Instances { 'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/NewsEye-Austrian-line.
NewsEye Austrian - line level
Table of Contents
- NewsEye Austrian - line level
- Table of Contents
- Dataset Description
- Languages
- Dataset Structure
- Data Instances
- Data Fields
Dataset Description
- Homepage: NewsEye project
- Source: Zenodo
- Point of Contact: TEKLIA
Dataset Summary
The dataset comprises Austrian newspaper pages from 19th and early 20th century. The images were provided by the Austrian National Library.
Languages
The documents are in Austrian German with the Fraktur font.
Note that all images are resized to a fixed height of 128 pixels.
Dataset Structure
Data Instances
{
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=4300x128 at 0x1A800E8E190,
'text': 'Mann; und als wir uns zum Angriff stark genug'
}Data Fields
image: a PIL.Image.Image object containing the image. Note that when accessing the image column (using dataset[0]["image"]), the image file is automatically decoded. Decoding of a large number of image files might take a significant amount of time. Thus it is important to first query the sample index before the "image" column, i.e. dataset[0]["image"] should always be preferred over dataset["image"][0].text: the label transcription of the image.
