CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /US-PD-Newspapers 🇺🇸 US Public Domain Newspapers 🇺🇸 US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library. With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining. Content As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.texttext-generation10M<n<100M50 likes6.9k downloads3y agoHugging Face02PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes6.4k downloads3y agoHugging Face03biglam /europeana_newspapers Dataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiative, this unofficial version extracts text content from ALTO XML and… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana_newspapers.imagetext-generation10M<n<100M63 likes1k downloads3d agoHugging Face04histde /ddb-newspaper-corpus 📰 DDB Newspaper Corpus A corpus of 11,551,703 pages of historical German newspapers in the public domain, harvested from the Deutsche Digitale Bibliothek (DDB) and its Zeitungsportal. It covers 1,607,744 issues from 796 newspapers published between 1638 and 1964, totalling 25.8 billion whitespace tokens (157.9 billion characters) of OCR fulltext. Every page carries an explicit per-page license (Public Domain Mark or CC0) and links back to its full-resolution scan (IIIF) and its… See the full description on the dataset page: https://huggingface.co/datasets/histde/ddb-newspaper-corpus.texttext-generation10M<n<100M12 likes860 downloads2mo agoHugging Face05PleIAs /NewZealand-PD-Newspapers New Zealand Public Domain Newspapers Dataset Card Dataset Overview Dataset Name: New Zealand Public Domain Newspapers Description: The New Zealand Public Domain Newspapers dataset comprises a collection of historical newspapers from New Zealand. The dataset is organized into Parquet files divided by year. Each file contains detailed information about the newspaper articles, including metadata extracted from XML files. Languages Covered: The dataset primarily contains… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/NewZealand-PD-Newspapers.texttext-generation1M<n<10M0 likes445 downloads2y agoHugging Face06community-datasets /id_newspapers_2018 Dataset Card for Indonesian Newspapers 2018 Dataset Summary The dataset contains around 500K articles (136M of words) from 7 Indonesian newspapers: Detik, Kompas, Tempo, CNN Indonesia, Sindo, Republika and Poskota. The articles are dated between 1st January 2018 and 20th August 2018 (with few exceptions dated earlier). The size of uncompressed 500K json files (newspapers-json.tgz) is around 2.2GB, and the cleaned uncompressed in a big text file (newspapers.txt.gz) is… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/id_newspapers_2018.texttext-generation100K<n<1M6 likes382 downloads2y agoHugging Face07biglam /hmd_newspapers Dataset Card for Heritage Made Digital Newspapers Dataset Summary This dataset contains text extracted at the article level from historic digitised newspapers from the Heritage Made Digital newspaper digitisation program at the British Library. The newspapers in the dataset were published between 1800 and 1896. This dataset contains ~2.5 billion tokens and 3,065,408 articles. The dataset contains text generated from Optical Character Recognition software on digitised… See the full description on the dataset page: https://huggingface.co/datasets/biglam/hmd_newspapers.tabulartext-generation1M<n<10M10 likes377 downloads3y agoHugging Face08storytracer /German-PD-Newspapers Dataset Card for Public Domain Newspapers (German) This dataset contains 13 billion words of OCR text extracted from German historical newspapers. Dataset Details Dataset Description Curated by: Sebastian Majstorovic Language(s) (NLP): German License: Dataset: CC0, Texts: Public Domain Dataset Sources [optional] Repository: https://www.deutsche-digitale-bibliothek.de/newspaper Copyright & License The newspapers texts have been… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/German-PD-Newspapers.texttext-generation1M<n<10M5 likes303 downloads3y agoHugging Face09bnl-data /bnl_newspapers Dataset Card for BnL Historical Newspapers Dataset Summary The BnL has digitised over 800.000 pages of Luxembourg newspapers. This dataset currently has one configuration covering a subset of these newspapers, which sit under the "Processed Datasets" collection. The BNL: processed all newspapers and monographs that are in the public domain and extracted the full text and associated meta data of every single article, section, advertisement… The result is a large number of… See the full description on the dataset page: https://huggingface.co/datasets/bnl-data/bnl_newspapers.texttext-generation100K<n<1M3 likes202 downloads3y agoHugging Face10biglam /bnl_newspapers1841-1879 Dataset Card for BnL Newspapers 1841-1881 Dataset Summary 592.192 articles from historical newspapers (1841-1881) along with metadata and the full text. 21 newspaper titles 24.415 newspaper issues 99.957 scanned pages Transcribed using a variety of OCR engines and corrected using https://github.com/natliblux/nautilusocr (95% threshold) Public Domain, CC0 (See copyright notice) The newspapers used are: Der Arbeiter (1878-1881) L'Arlequin (1848-1848) L'Avenir… See the full description on the dataset page: https://huggingface.co/datasets/biglam/bnl_newspapers1841-1879.texttext-generation100K<n<1M2 likes126 downloads2mo agoHugging Face11biglam /divergent-discourses-tibetan-newspapers Divergent Discourses — Early Tibetan Newspapers, 1950–1965 523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965, produced by the Divergent Discourses project (SOAS University of London and Leipzig University, with Trinity College Dublin). This is not a flat text dump. Each row is one text region from a scanned page, retaining its reading-order position, region type, source newspaper, and issue date — so page structure survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.tabulartext-generation100K<n<1M1 likes72 downloads2mo agoHugging Face12emdemor /news-of-the-brazilian-newspaper News of the Brazilian Newspaper This repository contains a comprehensive dataset of news articles from a Brazilian newspaper, Folha de São Paulo (http://www.folha.uol.com.br/). The dataset includes 167,053 examples of news articles, comprising headlines, URLs of articles, complete articles, and their respective categories. Dataset Creation The headlines were initially gathered from Inshorts and were then used to scrape the complete news articles from Folha de São Paulo.… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/news-of-the-brazilian-newspaper.texttext-classification100K<n<1M2 likes47 downloads2y agoHugging Face13GaloisField2718 /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/GaloisField2718/French-PD-Newspapers.tabulartext-generation1M<n<10M0 likes7 downloads7mo agoHugging Face14iknowyoursecret /US-PD-Newspapers 🇺🇸 US Public Domain Newspapers 🇺🇸 US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library. With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining. Content As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/iknowyoursecret/US-PD-Newspapers.texttext-generation10M<n<100M0 likes5 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.