newspaper
institutional-newspapers-bpl
📰 Institutional Newspapers: Boston Public Library
A structured dataset derived from the Boston Public Library's public domain
newspapers collection, produced by the Institutional Data
Initiative in collaboration with Boston Public Library.
1,473,635 public domain newspaper scans, published between 1795 and 1930
83,147,041 individual crops segmented from those scans
16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract
Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.19c_newspapers_images_altoUS-PD-Newspapers
🇺🇸 US Public Domain Newspapers 🇺🇸
US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library.
With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining.
Content
As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.French-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.newspaper-pagesnewspaper-ocr
