Teklia/PELLET-Casimir-Marius-line
PELLET Casimir Marius - Line level Dataset Summary The PELLET Casimir Marius dataset includes 100 annotated French letters written between 1914 and 1918. Annotations were done at line-level and all images do not have any text. Note that all images are resized to a fixed height of 128 pixels. Languages All the documents in the dataset are written in French. Dataset Structure Data Instances { 'image':… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/PELLET-Casimir-Marius-line.
PELLET Casimir Marius - Line level
Table of Contents
- PELLET Casimir Marius - Line level
- Table of Contents
- Dataset Description
- Dataset Summary
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Usage with the PyLaia library
Dataset Description
Dataset Summary
The PELLET Casimir Marius dataset includes 100 annotated French letters written between 1914 and 1918. Annotations were done at line-level and all images do not have any text.
Note that all images are resized to a fixed height of 128 pixels.
Languages
All the documents in the dataset are written in French.
Dataset Structure
Data Instances
{
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=1684x128 at 0x1A800E8E190,
'text': 'LE HAVRE - panorama de la rue de Paris'
}Data Fields
image: a PIL.Image.Image object containing the image. Note that when accessing the image column (using dataset[0]["image"]), the image file is automatically decoded. Decoding of a large number of image files might take a significant amount of time. Thus it is important to first query the sample index before the "image" column, i.e. dataset[0]["image"] should always be preferred over dataset["image"][0].text: the label transcription of the image.
Usage with the PyLaia library
- Clone the repository via
- the Settings on the UI,
- or
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/Teklia/PELLET-Casimir-Marius-line - The dataset is available in PyLaia format, in the
./pylaiafolder.
You can use this dataset to:
- train a new PyLaia model,
- assess your model's performance against this dataset.
