CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-index /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.texttext-generation10M<n<100M343 likes57k downloads1mo agoHugging Face02ruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B39 likes46k downloads7d agoHugging Face03PleIAs /US-PD-Newspapers 🇺🇸 US Public Domain Newspapers 🇺🇸 US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library. With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining. Content As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.texttext-generation10M<n<100M50 likes7.2k downloads3y agoHugging Face04PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes6.3k downloads3y agoHugging Face05dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face06openalphalab /gdelt-news GDELT News Reconstructions Content: multilingual news text reconstructed from GDELT Web News NGrams 3.0, including Type 1 and Type 2. Format: Zstandard-compressed Parquet only, with small manifests and a coverage checkpoint. Columns: date, language, source_url, text, observation_id, type, metadata. Metadata: source-file timestamp, source checksum and reconstruction diagnostics. Observation IDs are stable for the same source group. No country, publisher, author or external… See the full description on the dataset page: https://huggingface.co/datasets/openalphalab/gdelt-news.text-retrieval0 likes3.9k downloads0m agoHugging Face07vblagoje /cc_news Dataset Card for CC-News Dataset Summary CC-News dataset contains news articles from news sites all over the world. The data is available on AWS S3 in the Common Crawl bucket at /crawl-data/CC-NEWS/. This version of the dataset has been prepared using news-please - an integrated web crawler and information extractor for news.It contains 708241 English language news articles published between Jan 2017 and December 2019. It represents a small portion of the English… See the full description on the dataset page: https://huggingface.co/datasets/vblagoje/cc_news.imagetext-generation100K<n<1M70 likes3.4k downloads3y agoHugging Face08JulesBelveze /tldr_news Dataset Card for tldr_news Dataset Summary The tldr_news dataset was constructed by collecting daily tech newsletters from TLDR. For every piece of news, the title, content, category, section, and source URLs were extracted. The dataset now covers multiple TLDR newsletters including AI, tech, crypto, and other categories. This dataset can be used for various NLP tasks including: Headline generation Text summarization News categorization Content classification by… See the full description on the dataset page: https://huggingface.co/datasets/JulesBelveze/tldr_news.textsummarization10K<n<100K26 likes2.1k downloads9mo agoHugging Face09ctoraman /sozcu-news-2014(This dataset contains raw text, which are unlabeled.) 1,656 Turkish news articles from Sözcü Newspaper (http://www.sozcu.com.tr) between December 20, 2013, and March 11, 2014. GitHub Repo: https://github.com/BilkentInformationRetrievalGroup/TUBITAK113E249/ If you would like to use any material in this repository, please cite this paper: Toraman, C. and Can, F. (2017), Discovering story chains: A framework based on zigzagged search and news actors. Journal of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/sozcu-news-2014.texttext-generation1K<n<10K0 likes1.8k downloads3y agoHugging Face10BAAI /IndustryCorpus_news[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_news.texttext-generation100M<n<1B5 likes1.8k downloads1mo agoHugging Face11whiskey1983 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/whiskey1983/hacker-news.tabulartext-generation10M<n<100M0 likes1.6k downloads6mo agoHugging Face12MongoDB /tech-news-embeddings Overview HackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023. To further enhance the dataset's utility, a new embedding field and vector embedding for every datapoint have been added using the OpenAI EMBEDDING_MODEL = "text-embedding-3-small", with an EMBEDDING_DIMENSION of 256. Notably, this extension with vector embeddings only contains a portion of the original dataset, 1576528… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/tech-news-embeddings.textquestion-answering1M<n<10M6 likes1.6k downloads3y agoHugging Face13Whiteglove44 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/Whiteglove44/hacker-news.tabulartext-generation10M<n<100M0 likes1.6k downloads6mo agoHugging Face14Yahoo-Finance-News /FineWeb2024 FineWeb-Edu 2024 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2024. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2024 Rows 162,500,784… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2024.tabulartext-generation100M<n<1B0 likes1k downloads7d agoHugging Face15histde /ddb-newspaper-corpus 📰 DDB Newspaper Corpus A corpus of 11,551,703 pages of historical German newspapers in the public domain, harvested from the Deutsche Digitale Bibliothek (DDB) and its Zeitungsportal. It covers 1,607,744 issues from 796 newspapers published between 1638 and 1964, totalling 25.8 billion whitespace tokens (157.9 billion characters) of OCR fulltext. Every page carries an explicit per-page license (Public Domain Mark or CC0) and links back to its full-resolution scan (IIIF) and its… See the full description on the dataset page: https://huggingface.co/datasets/histde/ddb-newspaper-corpus.texttext-generation10M<n<100M12 likes896 downloads2mo agoHugging Face16biglam /europeana_newspapers Dataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiative, this unofficial version extracts text content from ALTO XML and… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana_newspapers.imagetext-generation10M<n<100M62 likes860 downloads1d agoHugging Face17naveenreddie-18 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/naveenreddie-18/hacker-news.tabulartext-generation10M<n<100M0 likes857 downloads6mo agoHugging Face18Yahoo-Finance-News /FineWeb2025 FineWeb-Edu 2025 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2025 Rows 99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.tabulartext-generation10M<n<100M1 likes745 downloads7d agoHugging Face19LeoCordoba /CC-NEWS-ES Dataset Card for CC-NEWS-ES Dataset Summary CC-NEWS-ES is a Spanish-language dataset of news. The corpus was generated by extracting the Spanish articles from CC-NEWS (news index of Common Crawl) of 2019. For doing that FastText model was used for language prediction. It contains a total of 7,473,286 texts and 1,812,009,283 words distributed as follows: domain texts words ar 532703 1.45127e+08 bo 29557 7.28996e+06 br 107 14207 cl 116661 3.34633e+07 co… See the full description on the dataset page: https://huggingface.co/datasets/LeoCordoba/CC-NEWS-ES.textsummarization1M<n<10M12 likes673 downloads4y agoHugging Face20nikhilambhure00 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/nikhilambhure00/hacker-news.tabulartext-generation10M<n<100M0 likes646 downloads6mo agoHugging Face21Scottswi /fenrix-financial-news-lake FENRIX Financial News Lake Topic-organized financial news data for portfolio-management decision research. The verified sanitized backing lake currently contains 37,571,131 rows across 1,461 Parquet shards and 8 sources. The product layer is organized by article meaning, not raw source lineage: article_group -> article_type -> year/month -> rows Source, license, URL, and provenance fields remain row-level audit metadata. Recommended Loading from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Scottswi/fenrix-financial-news-lake.text-classification10M<n<100M1 likes630 downloads3mo agoHugging Face22Yahoo-Finance-News /FineWeb-2023 FineWeb-Edu 2023 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2023. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2023 Rows 104,280,950… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb-2023.tabulartext-generation100M<n<1B0 likes620 downloads7d agoHugging Face23batubayk /TR-News Citation If you use the dataset, please cite the paper: @article{10.1007/s10579-021-09568-y, year = {2022}, title = {{Abstractive text summarization and new large-scale datasets for agglutinative languages Turkish and Hungarian}}, author = {Baykara, Batuhan and Güngör, Tunga}, journal = {Language Resources and Evaluation}, issn = {1574-020X}, doi = {10.1007/s10579-021-09568-y}, pages = {1--35}} textsummarization100K<n<1M19 likes609 downloads4y agoHugging Face24MurphyA /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/MurphyA/hacker-news.tabulartext-generation10M<n<100M0 likes525 downloads6mo agoHugging Face25PleIAs /NewZealand-PD-Newspapers New Zealand Public Domain Newspapers Dataset Card Dataset Overview Dataset Name: New Zealand Public Domain Newspapers Description: The New Zealand Public Domain Newspapers dataset comprises a collection of historical newspapers from New Zealand. The dataset is organized into Parquet files divided by year. Each file contains detailed information about the newspaper articles, including metadata extracted from XML files. Languages Covered: The dataset primarily contains… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/NewZealand-PD-Newspapers.texttext-generation1M<n<10M0 likes457 downloads2y agoHugging Face26maryamfakhari /crypto-news-coindesk-2020-2025 CoinDesk Cryptocurrency News Dataset (2020–2025) This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics. Time Period January 1, 2020 – January 1, 2025 Content Overview Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.imagetext-classification100K<n<1M2 likes447 downloads1mo agoHugging Face27vietgpt /binhvq_news_vi Binhvq News Source: https://github.com/binhvq/news-corpus Num examples: 19,365,593 Language: Vietnamese from datasets import load_dataset load_dataset("tdtunlp/binhvq_news_vi") texttext-generation10M<n<100M8 likes405 downloads3y agoHugging Face28biglam /hmd_newspapers Dataset Card for Heritage Made Digital Newspapers Dataset Summary This dataset contains text extracted at the article level from historic digitised newspapers from the Heritage Made Digital newspaper digitisation program at the British Library. The newspapers in the dataset were published between 1800 and 1896. This dataset contains ~2.5 billion tokens and 3,065,408 articles. The dataset contains text generated from Optical Character Recognition software on digitised… See the full description on the dataset page: https://huggingface.co/datasets/biglam/hmd_newspapers.tabulartext-generation1M<n<10M10 likes389 downloads3y agoHugging Face29community-datasets /id_newspapers_2018 Dataset Card for Indonesian Newspapers 2018 Dataset Summary The dataset contains around 500K articles (136M of words) from 7 Indonesian newspapers: Detik, Kompas, Tempo, CNN Indonesia, Sindo, Republika and Poskota. The articles are dated between 1st January 2018 and 20th August 2018 (with few exceptions dated earlier). The size of uncompressed 500K json files (newspapers-json.tgz) is around 2.2GB, and the cleaned uncompressed in a big text file (newspapers.txt.gz) is… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/id_newspapers_2018.texttext-generation100K<n<1M6 likes382 downloads2y agoHugging Face30eduagarcia /cc_news_pt_v2 Dataset Summary This version of the dataset is the portuguese subset from stanford-oval/ccnews. CC-News-PT v2 is a curation of +11 million news articles from CommonCrawl News in the Portuguese language, from the beginning (2016) to June of 2024. The data has been cleaned and deduplicated, and language of articles have been detected and filtered. The process is similar to what HuggingFace's DataTrove does. For license information, please refer to CommonCrawl's Terms of Use.… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/cc_news_pt_v2.imagetext-classification10M<n<100M4 likes382 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.