CoolFace
Datasetpublic

evndttr/Notes_de_lecture_PUL-Thevet_1575_Cosmographie

Notes de lecture, ca. 1578-1612 - Notes on André Thevet, Cosmographie universelle (1575) Description This dataset contains a excerpted transcription of Notes de lecture, ca. 1578-1612 — an anonymous commonplace book held at the Princeton University Library.This section of the manuscript records excerpts, paraphrases, and annotations taken from Thevet’s Cosmographie universelle (Paris, 1575). The present dataset includes a 15-page excerpt transcribed and aligned to… See the full description on the dataset page: https://huggingface.co/datasets/evndttr/Notes_de_lecture_PUL-Thevet_1575_Cosmographie.

sourceHugging Facecc-by-nc-4.0updated 11mo agoView on Hugging Face
0likes17downloads
Dataset Card

Notes de lecture, ca. 1578-1612 - Notes on André Thevet, Cosmographie universelle (1575)

Description

This dataset contains a excerpted transcription of Notes de lecture, ca. 1578-1612 — an anonymous commonplace book held at the Princeton University Library. This section of the manuscript records excerpts, paraphrases, and annotations taken from Thevet’s Cosmographie universelle (Paris, 1575).

The present dataset includes a 15-page excerpt transcribed and aligned to corresponding image segments as part of a 2023 Fellowship at the Princeton University Center for Digital Humanities (CDH).

Catalog record: Princeton University Library Catalog — 9960613933506421


Dataset contents

FieldTypeDescriptionExample
idstringUnique line identifier (page + line number)P001_L001
imagestringRelative path or URL to cropped line imageimages/P001_L003.jpg
pagestringSource page identifierP001
labelstringDiplomatic transcription of the handwritten lineCe qui ensuyt a este pris sur la cosmographie

Transcription parameters

  • Spelling, capitalization, and punctuation preserved as written.
  • Abbreviations are not expanded

The transcription was completed manually using Transkribus for segmentation and line-level annotation, exported in PAGE XML (2019-07-15) format, and later converted to tabular data for reuse and model training.


Use cases

This dataset can support:

  • Training or evaluating OCR/HTR models on early modern French handwriting
  • Text reuse and intertextuality studies (commonplace book ↔ Thevet’s printed text)