CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01institutional /institutional-newspapers-bplgated 📰 Institutional Newspapers: Boston Public Library A structured dataset derived from the Boston Public Library's public domain newspapers collection, produced by the Institutional Data Initiative in collaboration with Boston Public Library. 1,473,635 public domain newspaper scans, published between 1795 and 1930 83,147,041 individual crops segmented from those scans 16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.image1M<n<10M10 likes9.2k downloads1mo agoHugging Face02ambrosfitz /19c_newspapers_images_altotabular100K<n<1M4 likes8.2k downloads3mo agoHugging Face03PleIAs /US-PD-Newspapers 🇺🇸 US Public Domain Newspapers 🇺🇸 US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library. With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining. Content As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.texttext-generation10M<n<100M50 likes7.2k downloads3y agoHugging Face04PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes6.3k downloads3y agoHugging Face05PleIAs /Spanish-PD-Newspapers 🇪🇸 Spanish Public Domain Newspapers 🇪🇸 Spanish-Public Domain-Newspapers or Spanish-PD-Newspapers is a large collection aiming to aggregate all Spanish monographies in the public domain. As of March 2024, with Spanish-PD-Books, it is the biggest Spanish open corpus. Dataset summary The collection contains 247,491 individual texts making up 2,697,414,811 words recovered from multiple sources, including Spanish leading cultural heritage institution Biblioteca Digital… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Spanish-PD-Newspapers.text100K<n<1M4 likes909 downloads3y agoHugging Face06histde /ddb-newspaper-corpus 📰 DDB Newspaper Corpus A corpus of 11,551,703 pages of historical German newspapers in the public domain, harvested from the Deutsche Digitale Bibliothek (DDB) and its Zeitungsportal. It covers 1,607,744 issues from 796 newspapers published between 1638 and 1964, totalling 25.8 billion whitespace tokens (157.9 billion characters) of OCR fulltext. Every page carries an explicit per-page license (Public Domain Mark or CC0) and links back to its full-resolution scan (IIIF) and its… See the full description on the dataset page: https://huggingface.co/datasets/histde/ddb-newspaper-corpus.texttext-generation10M<n<100M12 likes896 downloads2mo agoHugging Face07biglam /europeana_newspapers Dataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiative, this unofficial version extracts text content from ALTO XML and… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana_newspapers.imagetext-generation10M<n<100M61 likes860 downloads19h agoHugging Face08PleIAs /NewZealand-PD-Newspapers New Zealand Public Domain Newspapers Dataset Card Dataset Overview Dataset Name: New Zealand Public Domain Newspapers Description: The New Zealand Public Domain Newspapers dataset comprises a collection of historical newspapers from New Zealand. The dataset is organized into Parquet files divided by year. Each file contains detailed information about the newspaper articles, including metadata extracted from XML files. Languages Covered: The dataset primarily contains… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/NewZealand-PD-Newspapers.texttext-generation1M<n<10M0 likes457 downloads2y agoHugging Face09biglam /hmd_newspapers Dataset Card for Heritage Made Digital Newspapers Dataset Summary This dataset contains text extracted at the article level from historic digitised newspapers from the Heritage Made Digital newspaper digitisation program at the British Library. The newspapers in the dataset were published between 1800 and 1896. This dataset contains ~2.5 billion tokens and 3,065,408 articles. The dataset contains text generated from Optical Character Recognition software on digitised… See the full description on the dataset page: https://huggingface.co/datasets/biglam/hmd_newspapers.tabulartext-generation1M<n<10M10 likes389 downloads3y agoHugging Face10community-datasets /id_newspapers_2018 Dataset Card for Indonesian Newspapers 2018 Dataset Summary The dataset contains around 500K articles (136M of words) from 7 Indonesian newspapers: Detik, Kompas, Tempo, CNN Indonesia, Sindo, Republika and Poskota. The articles are dated between 1st January 2018 and 20th August 2018 (with few exceptions dated earlier). The size of uncompressed 500K json files (newspapers-json.tgz) is around 2.2GB, and the cleaned uncompressed in a big text file (newspapers.txt.gz) is… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/id_newspapers_2018.texttext-generation100K<n<1M6 likes382 downloads2y agoHugging Face11davanstrien /newspaper_navigatorimageimage-to-text10M<n<100M0 likes304 downloads4y agoHugging Face12storytracer /German-PD-Newspapers Dataset Card for Public Domain Newspapers (German) This dataset contains 13 billion words of OCR text extracted from German historical newspapers. Dataset Details Dataset Description Curated by: Sebastian Majstorovic Language(s) (NLP): German License: Dataset: CC0, Texts: Public Domain Dataset Sources [optional] Repository: https://www.deutsche-digitale-bibliothek.de/newspaper Copyright & License The newspapers texts have been… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/German-PD-Newspapers.texttext-generation1M<n<10M5 likes301 downloads3y agoHugging Face13chcaa /eno-newspapers-enriched Danish Historical Newspaper Articles Dataset (enriched) This dataset contains approximately 4.9 million Danish historical newspaper articles (1666–1850) with document embeddings and assigned fictionality tags, providing a comprehensive resource for studying Danish language, culture, and history through primary journalistic sources. Dataset Details Dataset Description This dataset comprises digitized newspaper articles from Danish newspapers… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/eno-newspapers-enriched.tabular1M<n<10M1 likes283 downloads3mo agoHugging Face14biglam /newspaper-navigator Dataset Card for Newspaper Navigator Dataset Summary This dataset provides a Parquet-converted version of the Newspaper Navigator dataset from the Library of Congress. Originally released as JSON, Newspaper Navigator contains over 16 million pages of historic US newspapers annotated with bounding boxes, predicted visual types (e.g., photographs, maps), and OCR content. This work was carried out as part of a project by Benjamin Germain Lee et al. This version of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/newspaper-navigator.imageimage-classification1M<n<10M14 likes276 downloads1y agoHugging Face15bnl-data /bnl_newspapers Dataset Card for BnL Historical Newspapers Dataset Summary The BnL has digitised over 800.000 pages of Luxembourg newspapers. This dataset currently has one configuration covering a subset of these newspapers, which sit under the "Processed Datasets" collection. The BNL: processed all newspapers and monographs that are in the public domain and extracted the full text and associated meta data of every single article, section, advertisement… The result is a large number of… See the full description on the dataset page: https://huggingface.co/datasets/bnl-data/bnl_newspapers.texttext-generation100K<n<1M3 likes205 downloads3y agoHugging Face16davanstrien /newspapers-with-images Europeana Newspapers Sample Dataset Dataset Description This is a curated sample from the Europeana newspapers dataset, prepared for Vision-Language Model (VLM) experiments. Dataset Sources Original Dataset: biglam/europeana_newspapers Source: Europeana digital library License: See original dataset Dataset Structure Each sample contains a newspaper page image downloaded via IIIF along with associated metadata. Data… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/newspapers-with-images.image10K<n<100K1 likes203 downloads11mo agoHugging Face17Reubencf /konkani-newspapers Konkani Newspapers (Vauraddeancho Ixtt, 2012-2016) Scanned pages of Vauraddeancho Ixtt (Vauraddeancho Ixtt / वावराड्यांचो इष्ट, "Worker's Friend"), a Konkani-language weekly newspaper published from Pilar, Goa, India, and billed on its own masthead as "Goa's only Konkani Weekly since 1933". This is a page-image dataset, not a text corpus. Every row is one full scanned newspaper page, together with the date and page number parsed from the source filenames. There is no… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/konkani-newspapers.imageimage-to-text1K<n<10K0 likes135 downloads23d agoHugging Face18biglam /bnl_ground_truth_newspapers_before_1878 Dataset description 33.000 transcribed text lines from historical newspapers (before 1878) along with the cropped images of the original scans Text line based OCR 19.000 text lines in Antiqua 14.000 text lines in Fraktur Transcribed using double-keying (99.95% accuracy) Public Domain, CC0 (See copyright notice) Best for training an OCR engine The newspapers used are: Le Gratis luxembourgeois (1857-1858) Luxemburger Volks-Freund (1869-1876) L'Arlequin (1848-1848) Courrier du… See the full description on the dataset page: https://huggingface.co/datasets/biglam/bnl_ground_truth_newspapers_before_1878.image10K<n<100K2 likes134 downloads2mo agoHugging Face19zabir-nabil /bangla_newspaper_dataset Bangla Newspaper Dataset 400k+ bangla news samples, 25+ categories Source Data collected from https://www.prothomalo.com/archive [Copyright owned by the actual source] Github Github repository (Bi-LSTM Baseline): https://github.com/zabir-nabil/bangla-news-rnn Kaggle Version Kaggle Dataset: https://www.kaggle.com/datasets/furcifer/bangla-newspaper-dataset Inspiration The dataset can be used for Bangla text classification and generation… See the full description on the dataset page: https://huggingface.co/datasets/zabir-nabil/bangla_newspaper_dataset.tabulartext-classification100K<n<1M3 likes131 downloads2y agoHugging Face20biglam /bnl_newspapers1841-1879 Dataset Card for BnL Newspapers 1841-1881 Dataset Summary 592.192 articles from historical newspapers (1841-1881) along with metadata and the full text. 21 newspaper titles 24.415 newspaper issues 99.957 scanned pages Transcribed using a variety of OCR engines and corrected using https://github.com/natliblux/nautilusocr (95% threshold) Public Domain, CC0 (See copyright notice) The newspapers used are: Der Arbeiter (1878-1881) L'Arlequin (1848-1848) L'Avenir… See the full description on the dataset page: https://huggingface.co/datasets/biglam/bnl_newspapers1841-1879.texttext-generation100K<n<1M2 likes109 downloads2mo agoHugging Face21APProjects /us-media-newspaper-publishing-telecom-layoffs-warn-act-notices-daily US media, newspaper, publishing and telecom layoffs — the actual WARN Act filings, rebuilt every day Last rebuilt: 2026-09-24. 846 layoff and closure notices filed by newspapers and newspaper chains, broadcasters and TV station groups, film and game studios, magazine and book publishers, commercial printers, advertising and marketing agencies, and wireless, cable and telephone carriers and their call-centre contractors with US state labor departments — 101,859 workers, 265… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-media-newspaper-publishing-telecom-layoffs-warn-act-notices-daily.texttabular-classificationn<1K0 likes107 downloads1h agoHugging Face22SinclairSchneider /political_reference_newspaperstext100K<n<1M0 likes96 downloads6mo agoHugging Face23Teklia /Newspapers-finlam-La-Liberte Newspaper dataset: Finlam La Liberté Dataset Summary The Finlam La Liberté dataset includes 1500 issues from La Liberté, a French newspaper, from 1925 to 1928. Each issue contains multiple pages, with one image for each page resized to a fixed height of 2500 pixels. The dataset can be used to train end-to-end newspaper understanding models, with tasks including: Text zone detection and classification Reading order detection Article separation Split… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/Newspapers-finlam-La-Liberte.imageimage-to-text1K<n<10K3 likes95 downloads2y agoHugging Face24letrinhan /vn-provinces-newspaper-magazine-offices Vietnam newspaper and magazine offices Vietnam newspaper and magazine offices. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Hero (continued) Comparison Color key Files provinces (1071 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx regions (102 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-newspaper-magazine-offices.tabular1K<n<10K0 likes89 downloads4d agoHugging Face25Teklia /Newspapers-finlam Newspaper segmentation dataset: Finlam Dataset Summary The Finlam dataset includes 149 French newspapers from the 19th to 20th centuries. Each newspaper contains multiple pages. Page images are resized to a fixed height of 2000 pixels. Each page contains multiple zones, with different information such as polygon, text, class, and order. Split set images newspapers train 623 129 val 50 10 test 48 10 Languages Most newspapers in… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/Newspapers-finlam.imagen<1K6 likes82 downloads2y agoHugging Face26ULRs /Urdu-Newspaper-Benchmark Urdu Newspaper Benchmark This dataset follows the Hugging Face ImageFolder format with a metadata.csv file. Split train: 829 examples Columns HR Image: high-resolution image LR Image: low-resolution image Transcription: ground-truth Urdu text imageimage-to-textn<1K1 likes81 downloads6mo agoHugging Face27MatValSE /EE_Pre1944_NewspapersOCR data repository for Estonian pre 1944 newspapers text1K<n<10K0 likes77 downloads1y agoHugging Face28biglam /divergent-discourses-tibetan-newspapers Divergent Discourses — Early Tibetan Newspapers, 1950–1965 523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965, produced by the Divergent Discourses project (SOAS University of London and Leipzig University, with Trinity College Dublin). This is not a flat text dump. Each row is one text region from a scanned page, retaining its reading-order position, region type, source newspaper, and issue date — so page structure survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.tabulartext-generation100K<n<1M1 likes72 downloads2mo agoHugging Face29davanstrien /newspapers-with-images-after-photography-big Europeana Newspapers Sample Dataset Dataset Description This is a curated sample from the Europeana newspapers dataset, prepared for Vision-Language Model (VLM) experiments. Dataset Sources Original Dataset: biglam/europeana_newspapers Source: Europeana digital library License: See original dataset Dataset Structure Each sample contains a newspaper page image downloaded via IIIF along with associated metadata. Data… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/newspapers-with-images-after-photography-big.image1K<n<10K0 likes68 downloads10mo agoHugging Face30biglam /europeana-newspapers-ground-truth Europeana Newspapers — Historical Newspapers Ground Truth 50 pages of digitised historical German newspapers from the Berlin State Library (Staatsbibliothek zu Berlin), with PAGE XML ground truth produced for the EU Europeana Newspapers project. Each row pairs three things: the page image, the human-corrected ground truth (regions, polygons, reading order, text), and the ALTO OCR output that ABBYY FineReader actually produced. That last column is what makes this an OCR… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana-newspapers-ground-truth.imageimage-segmentationn<1K0 likes56 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.