datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
newspaper-navigator
Dataset Card for Newspaper Navigator
Dataset Summary
This dataset provides a Parquet-converted version of the Newspaper Navigator dataset from the Library of Congress. Originally released as JSON, Newspaper Navigator contains over 16 million pages of historic US newspapers annotated with bounding boxes, predicted visual types (e.g., photographs, maps), and OCR content. This work was carried out as part of a project by Benjamin Germain Lee et al.
This version of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/newspaper-navigator.Newspapers-finlam-La-Liberte
Newspaper dataset: Finlam La Liberté
Dataset Summary
The Finlam La Liberté dataset includes 1500 issues from La Liberté, a French newspaper, from 1925 to 1928.
Each issue contains multiple pages, with one image for each page resized to a fixed height of 2500 pixels.
The dataset can be used to train end-to-end newspaper understanding models, with tasks including:
Text zone detection and classification
Reading order detection
Article separation
Split… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/Newspapers-finlam-La-Liberte.index-cards-peabody-newspaper
Peabody Newspaper Index Cards (Peabody Institute Library, MA)
3,694 typewritten index cards from the Peabody Institute Library — Sutton
Room Local History Resource Center (Peabody, Massachusetts), indexing people,
events, and news in South Danvers / Peabody as recorded in local newspapers.
The information was typed onto cards over decades by library staff as the local
newspaper-of-record archive's principal finding aid.
Plus a companion "Poor Family" genealogy index from the same… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-peabody-newspaper.
