datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
institutional-newspapers-bpl
📰 Institutional Newspapers: Boston Public Library
A structured dataset derived from the Boston Public Library's public domain
newspapers collection, produced by the Institutional Data
Initiative in collaboration with Boston Public Library.
1,473,635 public domain newspaper scans, published between 1795 and 1930
83,147,041 individual crops segmented from those scans
16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract
Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.19c_newspapers_images_altoUS-PD-Newspapers
🇺🇸 US Public Domain Newspapers 🇺🇸
US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library.
With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining.
Content
As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.French-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.Spanish-PD-Newspapers
🇪🇸 Spanish Public Domain Newspapers 🇪🇸
Spanish-Public Domain-Newspapers or Spanish-PD-Newspapers is a large collection aiming to aggregate all Spanish monographies in the public domain. As of March 2024, with Spanish-PD-Books, it is the biggest Spanish open corpus.
Dataset summary
The collection contains 247,491 individual texts making up 2,697,414,811 words recovered from multiple sources, including Spanish leading cultural heritage institution Biblioteca Digital… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Spanish-PD-Newspapers.ddb-newspaper-corpus
📰 DDB Newspaper Corpus
A corpus of 11,551,703 pages of historical German newspapers in the public domain, harvested from the Deutsche Digitale Bibliothek (DDB) and its Zeitungsportal.
It covers 1,607,744 issues from 796 newspapers published between 1638 and 1964, totalling 25.8 billion whitespace tokens (157.9 billion characters) of OCR fulltext. Every page carries an explicit per-page license (Public Domain Mark or CC0) and links back to its full-resolution scan (IIIF) and its… See the full description on the dataset page: https://huggingface.co/datasets/histde/ddb-newspaper-corpus.europeana_newspapers
Dataset Card for Europeana Newspapers
Dataset Overview
This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century.
Created by the BigLAM initiative, this unofficial version extracts text content from ALTO XML and… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana_newspapers.NewZealand-PD-Newspapers
New Zealand Public Domain Newspapers Dataset Card
Dataset Overview
Dataset Name: New Zealand Public Domain Newspapers
Description:
The New Zealand Public Domain Newspapers dataset comprises a collection of historical newspapers from New Zealand. The dataset is organized into Parquet files divided by year. Each file contains detailed information about the newspaper articles, including metadata extracted from XML files.
Languages Covered:
The dataset primarily contains… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/NewZealand-PD-Newspapers.hmd_newspapers
Dataset Card for Heritage Made Digital Newspapers
Dataset Summary
This dataset contains text extracted at the article level from historic digitised newspapers from the Heritage Made Digital newspaper digitisation program at the British Library. The newspapers in the dataset were published between 1800 and 1896. This dataset contains ~2.5 billion tokens and 3,065,408 articles.
The dataset contains text generated from Optical Character Recognition software on digitised… See the full description on the dataset page: https://huggingface.co/datasets/biglam/hmd_newspapers.id_newspapers_2018
Dataset Card for Indonesian Newspapers 2018
Dataset Summary
The dataset contains around 500K articles (136M of words) from 7 Indonesian newspapers: Detik, Kompas, Tempo,
CNN Indonesia, Sindo, Republika and Poskota. The articles are dated between 1st January 2018 and 20th August 2018
(with few exceptions dated earlier). The size of uncompressed 500K json files (newspapers-json.tgz) is around 2.2GB,
and the cleaned uncompressed in a big text file (newspapers.txt.gz) is… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/id_newspapers_2018.newspaper_navigatorGerman-PD-Newspapers
Dataset Card for Public Domain Newspapers (German)
This dataset contains 13 billion words of OCR text extracted from German historical newspapers.
Dataset Details
Dataset Description
Curated by: Sebastian Majstorovic
Language(s) (NLP): German
License: Dataset: CC0, Texts: Public Domain
Dataset Sources [optional]
Repository: https://www.deutsche-digitale-bibliothek.de/newspaper
Copyright & License
The newspapers texts have been… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/German-PD-Newspapers.eno-newspapers-enriched
Danish Historical Newspaper Articles Dataset (enriched)
This dataset contains approximately 4.9 million Danish historical newspaper articles (1666–1850) with document embeddings and assigned fictionality tags, providing a comprehensive resource for studying Danish language, culture, and history through primary journalistic sources.
Dataset Details
Dataset Description
This dataset comprises digitized newspaper articles from Danish newspapers… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/eno-newspapers-enriched.newspaper-navigator
Dataset Card for Newspaper Navigator
Dataset Summary
This dataset provides a Parquet-converted version of the Newspaper Navigator dataset from the Library of Congress. Originally released as JSON, Newspaper Navigator contains over 16 million pages of historic US newspapers annotated with bounding boxes, predicted visual types (e.g., photographs, maps), and OCR content. This work was carried out as part of a project by Benjamin Germain Lee et al.
This version of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/newspaper-navigator.bnl_newspapers
Dataset Card for BnL Historical Newspapers
Dataset Summary
The BnL has digitised over 800.000 pages of Luxembourg newspapers. This dataset currently has one configuration covering a subset of these newspapers, which sit under the "Processed Datasets" collection. The BNL:
processed all newspapers and monographs that are in the public domain and extracted the full text and associated meta data of every single article, section, advertisement… The result is a large number of… See the full description on the dataset page: https://huggingface.co/datasets/bnl-data/bnl_newspapers.newspapers-with-images
Europeana Newspapers Sample Dataset
Dataset Description
This is a curated sample from the Europeana newspapers dataset, prepared for Vision-Language Model (VLM) experiments.
Dataset Sources
Original Dataset: biglam/europeana_newspapers
Source: Europeana digital library
License: See original dataset
Dataset Structure
Each sample contains a newspaper page image downloaded via IIIF along with associated metadata.
Data… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/newspapers-with-images.konkani-newspapers
Konkani Newspapers (Vauraddeancho Ixtt, 2012-2016)
Scanned pages of Vauraddeancho Ixtt (Vauraddeancho Ixtt / वावराड्यांचो इष्ट,
"Worker's Friend"), a Konkani-language weekly newspaper published from Pilar, Goa,
India, and billed on its own masthead as "Goa's only Konkani Weekly since 1933".
This is a page-image dataset, not a text corpus. Every row is one full scanned
newspaper page, together with the date and page number parsed from the source
filenames. There is no… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/konkani-newspapers.bnl_ground_truth_newspapers_before_1878
Dataset description
33.000 transcribed text lines from historical newspapers (before 1878) along with the cropped images of the original scans
Text line based OCR
19.000 text lines in Antiqua
14.000 text lines in Fraktur
Transcribed using double-keying (99.95% accuracy)
Public Domain, CC0 (See copyright notice)
Best for training an OCR engine
The newspapers used are:
Le Gratis luxembourgeois (1857-1858)
Luxemburger Volks-Freund (1869-1876)
L'Arlequin (1848-1848)
Courrier du… See the full description on the dataset page: https://huggingface.co/datasets/biglam/bnl_ground_truth_newspapers_before_1878.bangla_newspaper_dataset
Bangla Newspaper Dataset
400k+ bangla news samples, 25+ categories
Source
Data collected from https://www.prothomalo.com/archive [Copyright owned by the actual source]
Github
Github repository (Bi-LSTM Baseline): https://github.com/zabir-nabil/bangla-news-rnn
Kaggle Version
Kaggle Dataset: https://www.kaggle.com/datasets/furcifer/bangla-newspaper-dataset
Inspiration
The dataset can be used for Bangla text classification and generation… See the full description on the dataset page: https://huggingface.co/datasets/zabir-nabil/bangla_newspaper_dataset.bnl_newspapers1841-1879
Dataset Card for BnL Newspapers 1841-1881
Dataset Summary
592.192 articles from historical newspapers (1841-1881) along with metadata and the full text.
21 newspaper titles
24.415 newspaper issues
99.957 scanned pages
Transcribed using a variety of OCR engines and corrected using https://github.com/natliblux/nautilusocr (95% threshold)
Public Domain, CC0 (See copyright notice)
The newspapers used are:
Der Arbeiter (1878-1881)
L'Arlequin (1848-1848)
L'Avenir… See the full description on the dataset page: https://huggingface.co/datasets/biglam/bnl_newspapers1841-1879.us-media-newspaper-publishing-telecom-layoffs-warn-act-notices-daily
US media, newspaper, publishing and telecom layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-24. 846 layoff and closure notices filed by
newspapers and newspaper chains, broadcasters and TV station groups, film and game studios, magazine and book publishers, commercial printers, advertising and marketing agencies, and wireless, cable and telephone carriers and their call-centre contractors with US state labor departments — 101,859 workers,
265… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-media-newspaper-publishing-telecom-layoffs-warn-act-notices-daily.political_reference_newspapersNewspapers-finlam-La-Liberte
Newspaper dataset: Finlam La Liberté
Dataset Summary
The Finlam La Liberté dataset includes 1500 issues from La Liberté, a French newspaper, from 1925 to 1928.
Each issue contains multiple pages, with one image for each page resized to a fixed height of 2500 pixels.
The dataset can be used to train end-to-end newspaper understanding models, with tasks including:
Text zone detection and classification
Reading order detection
Article separation
Split… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/Newspapers-finlam-La-Liberte.vn-provinces-newspaper-magazine-offices
Vietnam newspaper and magazine offices
Vietnam newspaper and magazine offices. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1071 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (102 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-newspaper-magazine-offices.Newspapers-finlam
Newspaper segmentation dataset: Finlam
Dataset Summary
The Finlam dataset includes 149 French newspapers from the 19th to 20th centuries.
Each newspaper contains multiple pages. Page images are resized to a fixed height of 2000 pixels.
Each page contains multiple zones, with different information such as polygon, text, class, and order.
Split
set
images
newspapers
train
623
129
val
50
10
test
48
10
Languages
Most newspapers in… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/Newspapers-finlam.Urdu-Newspaper-Benchmark
Urdu Newspaper Benchmark
This dataset follows the Hugging Face ImageFolder format with a metadata.csv file.
Split
train: 829 examples
Columns
HR Image: high-resolution image
LR Image: low-resolution image
Transcription: ground-truth Urdu text
EE_Pre1944_NewspapersOCR data repository for Estonian pre 1944 newspapers
divergent-discourses-tibetan-newspapers
Divergent Discourses — Early Tibetan Newspapers, 1950–1965
523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965,
produced by the Divergent Discourses project (SOAS University of London and Leipzig
University, with Trinity College Dublin).
This is not a flat text dump. Each row is one text region from a scanned page, retaining
its reading-order position, region type, source newspaper, and issue date — so page structure
survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.newspapers-with-images-after-photography-big
Europeana Newspapers Sample Dataset
Dataset Description
This is a curated sample from the Europeana newspapers dataset, prepared for Vision-Language Model (VLM) experiments.
Dataset Sources
Original Dataset: biglam/europeana_newspapers
Source: Europeana digital library
License: See original dataset
Dataset Structure
Each sample contains a newspaper page image downloaded via IIIF along with associated metadata.
Data… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/newspapers-with-images-after-photography-big.europeana-newspapers-ground-truth
Europeana Newspapers — Historical Newspapers Ground Truth
50 pages of digitised historical German newspapers from the Berlin State Library
(Staatsbibliothek zu Berlin), with PAGE XML ground truth produced for the EU
Europeana Newspapers project.
Each row pairs three things: the page image, the human-corrected ground truth
(regions, polygons, reading order, text), and the ALTO OCR output that ABBYY FineReader
actually produced. That last column is what makes this an OCR… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana-newspapers-ground-truth.
