datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
US-PD-Newspapers
🇺🇸 US Public Domain Newspapers 🇺🇸
US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library.
With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining.
Content
As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.French-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.europeana_newspapers
Dataset Card for Europeana Newspapers
Dataset Overview
This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century.
Created by the BigLAM initiative, this unofficial version extracts text content from ALTO XML and… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana_newspapers.ddb-newspaper-corpus
📰 DDB Newspaper Corpus
A corpus of 11,551,703 pages of historical German newspapers in the public domain, harvested from the Deutsche Digitale Bibliothek (DDB) and its Zeitungsportal.
It covers 1,607,744 issues from 796 newspapers published between 1638 and 1964, totalling 25.8 billion whitespace tokens (157.9 billion characters) of OCR fulltext. Every page carries an explicit per-page license (Public Domain Mark or CC0) and links back to its full-resolution scan (IIIF) and its… See the full description on the dataset page: https://huggingface.co/datasets/histde/ddb-newspaper-corpus.NewZealand-PD-Newspapers
New Zealand Public Domain Newspapers Dataset Card
Dataset Overview
Dataset Name: New Zealand Public Domain Newspapers
Description:
The New Zealand Public Domain Newspapers dataset comprises a collection of historical newspapers from New Zealand. The dataset is organized into Parquet files divided by year. Each file contains detailed information about the newspaper articles, including metadata extracted from XML files.
Languages Covered:
The dataset primarily contains… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/NewZealand-PD-Newspapers.id_newspapers_2018
Dataset Card for Indonesian Newspapers 2018
Dataset Summary
The dataset contains around 500K articles (136M of words) from 7 Indonesian newspapers: Detik, Kompas, Tempo,
CNN Indonesia, Sindo, Republika and Poskota. The articles are dated between 1st January 2018 and 20th August 2018
(with few exceptions dated earlier). The size of uncompressed 500K json files (newspapers-json.tgz) is around 2.2GB,
and the cleaned uncompressed in a big text file (newspapers.txt.gz) is… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/id_newspapers_2018.hmd_newspapers
Dataset Card for Heritage Made Digital Newspapers
Dataset Summary
This dataset contains text extracted at the article level from historic digitised newspapers from the Heritage Made Digital newspaper digitisation program at the British Library. The newspapers in the dataset were published between 1800 and 1896. This dataset contains ~2.5 billion tokens and 3,065,408 articles.
The dataset contains text generated from Optical Character Recognition software on digitised… See the full description on the dataset page: https://huggingface.co/datasets/biglam/hmd_newspapers.German-PD-Newspapers
Dataset Card for Public Domain Newspapers (German)
This dataset contains 13 billion words of OCR text extracted from German historical newspapers.
Dataset Details
Dataset Description
Curated by: Sebastian Majstorovic
Language(s) (NLP): German
License: Dataset: CC0, Texts: Public Domain
Dataset Sources [optional]
Repository: https://www.deutsche-digitale-bibliothek.de/newspaper
Copyright & License
The newspapers texts have been… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/German-PD-Newspapers.bnl_newspapers
Dataset Card for BnL Historical Newspapers
Dataset Summary
The BnL has digitised over 800.000 pages of Luxembourg newspapers. This dataset currently has one configuration covering a subset of these newspapers, which sit under the "Processed Datasets" collection. The BNL:
processed all newspapers and monographs that are in the public domain and extracted the full text and associated meta data of every single article, section, advertisement… The result is a large number of… See the full description on the dataset page: https://huggingface.co/datasets/bnl-data/bnl_newspapers.bnl_newspapers1841-1879
Dataset Card for BnL Newspapers 1841-1881
Dataset Summary
592.192 articles from historical newspapers (1841-1881) along with metadata and the full text.
21 newspaper titles
24.415 newspaper issues
99.957 scanned pages
Transcribed using a variety of OCR engines and corrected using https://github.com/natliblux/nautilusocr (95% threshold)
Public Domain, CC0 (See copyright notice)
The newspapers used are:
Der Arbeiter (1878-1881)
L'Arlequin (1848-1848)
L'Avenir… See the full description on the dataset page: https://huggingface.co/datasets/biglam/bnl_newspapers1841-1879.divergent-discourses-tibetan-newspapers
Divergent Discourses — Early Tibetan Newspapers, 1950–1965
523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965,
produced by the Divergent Discourses project (SOAS University of London and Leipzig
University, with Trinity College Dublin).
This is not a flat text dump. Each row is one text region from a scanned page, retaining
its reading-order position, region type, source newspaper, and issue date — so page structure
survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.news-of-the-brazilian-newspaper
News of the Brazilian Newspaper
This repository contains a comprehensive dataset of news articles from a Brazilian newspaper, Folha de São Paulo (http://www.folha.uol.com.br/). The dataset includes 167,053 examples of news articles, comprising headlines, URLs of articles, complete articles, and their respective categories.
Dataset Creation
The headlines were initially gathered from Inshorts and were then used to scrape the complete news articles from Folha de São Paulo.… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/news-of-the-brazilian-newspaper.French-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/GaloisField2718/French-PD-Newspapers.US-PD-Newspapers
🇺🇸 US Public Domain Newspapers 🇺🇸
US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library.
With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining.
Content
As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/iknowyoursecret/US-PD-Newspapers.
