CoolFace
Datasetpublic

Teklia/NewsEye-Austrian-line

NewsEye Austrian - line level Dataset Summary The dataset comprises Austrian newspaper pages from 19th and early 20th century. The images were provided by the Austrian National Library. Languages The documents are in Austrian German with the Fraktur font. Note that all images are resized to a fixed height of 128 pixels. Dataset Structure Data Instances { 'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/NewsEye-Austrian-line.

sourceHugging Facemitupdated 3y agoView on Hugging Face
4likes60downloads
Dataset Card

NewsEye Austrian - line level

Table of Contents

Dataset Description

Dataset Summary

The dataset comprises Austrian newspaper pages from 19th and early 20th century. The images were provided by the Austrian National Library.

Languages

The documents are in Austrian German with the Fraktur font.

Note that all images are resized to a fixed height of 128 pixels.

Dataset Structure

Data Instances

{
  'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=4300x128 at 0x1A800E8E190,
  'text': 'Mann; und als wir uns zum Angriff stark genug'
}

Data Fields

  • —image: a PIL.Image.Image object containing the image. Note that when accessing the image column (using dataset[0]["image"]), the image file is automatically decoded. Decoding of a large number of image files might take a significant amount of time. Thus it is important to first query the sample index before the "image" column, i.e. dataset[0]["image"] should always be preferred over dataset["image"][0].
  • —text: the label transcription of the image.