CoolFace
Datasetpublic

CATMuS/medieval-segmentation

Dataset Card for CATMuS Medieval (Segmentation Version) Join our Discord to ask questions about the dataset: Dataset Details CATMuS Medieval Segmentation (Consistent Approaches to Transcribing Manuscripts) is a specialized dataset designed for layout analysis of medieval manuscripts using the SegmOnto vocabulary for region and line classification. This dataset addresses the challenges associated with establishing consistent ground truth in layout analysis tasks… See the full description on the dataset page: https://huggingface.co/datasets/CATMuS/medieval-segmentation.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
7likes1.7kdownloads
Dataset Card

Dataset Card for CATMuS Medieval (Segmentation Version)

[image]

Join our Discord to ask questions about the dataset: ![Join the Discord](https://discord.gg/J38xgNEsGk)

Dataset Details

CATMuS Medieval Segmentation (Consistent Approaches to Transcribing Manuscripts) is a specialized dataset designed for layout analysis of medieval manuscripts using the SegmOnto vocabulary for region and line classification. This dataset addresses the challenges associated with establishing consistent ground truth in layout analysis tasks, particularly for the complex and heterogeneous historical sources of medieval manuscripts in Latin scripts from the 8th to the 15th century CE. It is a subset of the manuscript present in the CATMuS Medieval dataset, which focuses on HTR only.

The CATMuS dataset for layout analysis provides:

  • A uniform framework for annotation practices for the layout of medieval manuscripts.
  • A benchmarking environment for evaluating automatic layout analysis models across multiple dimensions thanks to some metadata (for now, century of production).
  • A benchmarking environment for other tasks (such as datation approaches).
  • A platform for exploratory work in computer vision and digital paleography focused on layout-based tasks, such as layout generation.

Developed through collaboration among various institutions and projects, CATMuS Medieval offers an inter-compatible dataset that spans over 200 manuscripts and incunabula in 10 different languages, containing a wealth of structural annotations using the SegmOnto vocabulary.

By ensuring consistency in layout analysis approaches, CATMuS aims to mitigate challenges arising from the diversity in standards for medieval manuscript analysis. It provides a comprehensive benchmark for evaluating layout analysis models on historical sources, facilitating advancements in the field of digital humanities.

Dataset Description

<!-- Provide a longer summary of what this dataset is. -->

  • Curated by: Thibault Clérice (Inria)
  • Funded by: BnF Datalab, Biblissima +, DIM PAMIR
  • License: CC-BY 4.0
Documents
traindevtestTotal
images13361911781705
manuscripts1592028207
Century coverage

As the number of images in each split. Images can represent two pages.

traindevtestTotal
Century:082002
Century:0911110112
Century:101103849
Century:11270027
Century:1219171046
Century:13230920259
Century:1424111139391
Century:155633619618
Century:161321752201
Lines
traindevtestTotal
Line:DefaultLine817831355412595107932
Line:DropCapitalLine11751051001380
Line:HeadingLine13817011652247
Line:InterlinearLine28082722345069
Line:MusicLine16700167
Line:TironianSignLine28200282
Zones
traindevtestTotal
Zone:DamageZone121013
Zone:DigitizationArtefactZone280028
Zone:DropCapitalZone15671021321801
Zone:GraphicZone300715322
Zone:MainZone23173652942976
Zone:MarginTextZone9161461991261
Zone:MusicZone17900179
Zone:NumberingZone63210295829
Zone:QuireMarksZone86915110
Zone:RunningTitleZone3409118449
Zone:SealZone3003
Zone:StampZone395549
Zone:TitlePageZone4127

Uses

Direct Use

  • Layout Analysis

Out-of-Scope Use

  • Text-To-Image

Dataset Structure

  • data contains 3 splits, which are loaded through load_dataset("CATMuS/medieval-segmentation"). They are the same split as Catmus Medieval (for HTR)
  • Each image is annotated with a
  • file_name (path from root)
  • shelfmark identifier
  • century datation information
  • project that originally produced the data
  • width of the page (in pixels)
  • height of the page (in pixels)
  • objects which contains sequences of values for each object found in the page:
  • id that are mainly used for parent relationship between blocks (such as columns) and lines
  • bbox (Shape: [x1, y1, x2, y2], top left to bottom right)
  • polygons (Shape: [x, y, x, y, x, y, ...])
  • category as a list of string using the first level of SegmOnto guidelines
  • type which is either block (area) or line.
  • parent which hold the id of the parent (is null for blocks, is nullable for line)

<!--

Dataset Creation

Curation Rationale

Motivation for the creation of this dataset.

[More Information Needed]

Source Data

This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...).

Data Collection and Processi
Who are the source data producers?

-->

Annotations

Annotation process

The annotation process is described in the dataset paper.

Who are the annotators?
  • Pinche, Ariane
  • Clérice, Thibault
  • Chagué, Alix
  • Camps, Jean-Baptiste
  • Vlachou-Efstathiou, Malamatenia
  • Gille Levenson, Matthias
  • Brisville-Fertin, Olivier
  • Boschetti, Federico
  • Fischer, Franz
  • Gervers, Michael
  • Boutreux, Agnès
  • Manton, Avery
  • Gabay, Simon
  • Bordier, Julie
  • Glaise, Anthony
  • Alba, Rachele
  • Rubin, Giorgia
  • White, Nick
  • Karaisl, Antonia
  • Leroy, Noé
  • Maulu, Marco
  • Biay, Sébastien
  • Cappe, Zoé
  • Konstantinova, Kristina
  • Boby, Victor
  • Christensen, Kelly
  • Pierreville, Corinne
  • Aruta, Davide
  • Lenzi, Martina
  • Le Huëron, Armelle
  • Possamaï, Marylène
  • Duval, Frédéric
  • Mariotti, Violetta
  • Morreale, Laura
  • Nolibois, Alice
  • Foehr-Janssens, Yasmina
  • Deleville, Prunelle
  • Carnaille, Camille
  • Lecomte, Sophie
  • Meylan, Aminoel
  • Ventura, Simone
  • Dugaz, Lucien
Software

The software to generate this version of the dataset was built by Thibault Clérice and William Mattingly

Bias, Risks, and Limitations

The data are skewed toward Old French, Middle Dutch and Spanish, specifically from the 14th century.

The only language that is represented over all centuries is Latin, and in each scripts. The other language with a coverage close to Latin is Old French.

Only one document is available in Old English.

Citation

BibTeX:

tex
@unpublished{clerice:hal-04453952,
  TITLE = {{CATMuS Medieval: A multilingual large-scale cross-century dataset in Latin script for handwritten text recognition and beyond}},
  AUTHOR = {Cl{\'e}rice, Thibault and Pinche, Ariane and Vlachou-Efstathiou, Malamatenia and Chagu{\'e}, Alix and Camps, Jean-Baptiste and Gille-Levenson, Matthias and Brisville-Fertin, Olivier and Fischer, Franz and Gervers, Michaels and Boutreux, Agn{\`e}s and Manton, Avery and Gabay, Simon and O'Connor, Patricia and Haverals, Wouter and Kestemont, Mike and Vandyck, Caroline and Kiessling, Benjamin},
  URL = {https://inria.hal.science/hal-04453952},
  NOTE = {working paper or preprint},
  YEAR = {2024},
  MONTH = Feb,
  KEYWORDS = {Historical sources ; medieval manuscripts ; Latin scripts ; benchmarking dataset ; multilingual ; handwritten text recognition},
  PDF = {https://inria.hal.science/hal-04453952/file/ICDAR24___CATMUS_Medieval-1.pdf},
  HAL_ID = {hal-04453952},
  HAL_VERSION = {v1},
}

APA:

Thibault Clérice, Ariane Pinche, Malamatenia Vlachou-Efstathiou, Alix Chagué, Jean-Baptiste Camps, et al.. CATMuS Medieval: A multilingual large-scale cross-century dataset in Latin script for handwritten text recognition and beyond. 2024. ⟨hal-04453952⟩

Dataset Card Contact

Thibault Clérice (first.last@inria.fr)