CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fancyzhx /ag_news Dataset Card for "ag_news" Dataset Summary AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc)… See the full description on the dataset page: https://huggingface.co/datasets/fancyzhx/ag_news.texttext-classification100K<n<1M195 likes88k downloads3y agoHugging Face02open-index /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.texttext-generation10M<n<100M343 likes47k downloads1mo agoHugging Face03ruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B39 likes46k downloads5d agoHugging Face04shash42 /forecast-news Forecast News Deduplicated daily news corpus used by forecast-sim and future-sim. Snapshot 31,859,020 articles 3,463 daily partitions Coverage: 2016-08-26 through 2026-08-31 Snapshot published: 2026-09-18 Stored data size: approximately 158.5 GiB The Parquet files are the canonical complete representation. The repository also contains daily JSONL files where available and compact headline JSON files used by article-browsing workflows. Layout Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.text10M<n<100M2 likes40k downloads4d agoHugging Face05aisbergpublicorganization /telegram-news-ua-dataset Aisberg Telegram News UA A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.texttext-classification100K<n<1M4 likes16k downloads1h agoHugging Face06AlphaDojo /dojo_stock_news Languages: 简体中文 · English dojo_stock_news — Stock News Overview Financial news linked to individual stocks: headline, summary, source, publish time, and URL. Files File Description data.parquet Full news archive Key Fields Field Description symbol Associated stock symbol (primary query key) title Headline description Summary body publish_date Publish date (YYYY-MM-DD or locale-specific text)… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_stock_news.text1M<n<10M0 likes14k downloads13h agoHugging Face07alexfabbri /multi_newsMulti-News, consists of news articles and human-written summaries of these articles from the site newser.com. Each summary is professionally written by editors and includes links to the original articles cited. There are two features: - document: text of news articles seperated by special token "|||||". - summary: news summary.summarization10K<n<100K78 likes11k downloads3y agoHugging Face08shash42 /forecast-news-embeddings Forecast News Embeddings Precomputed LanceDB table used for keyword, semantic, and hybrid retrieval in forecast-sim and future-sim. Snapshot 7,911,857 indexed source articles 16,207,764 text chunks Coverage: 2023-01-11 through 2026-08-31 Snapshot published: 2026-09-18 Lance dataset version: 856 Total artifact size: approximately 303.2 GiB Articles with empty searchable text are not represented. Long articles can produce multiple chunks, so the chunk count is… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news-embeddings.2 likes9.8k downloads4d agoHugging Face09SetFit /20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs: The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date. We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.text10K<n<100K21 likes9.4k downloads5y agoHugging Face10institutional /institutional-newspapers-bplgated 📰 Institutional Newspapers: Boston Public Library A structured dataset derived from the Boston Public Library's public domain newspapers collection, produced by the Institutional Data Initiative in collaboration with Boston Public Library. 1,473,635 public domain newspaper scans, published between 1795 and 1930 83,147,041 individual crops segmented from those scans 16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.image1M<n<10M9 likes9.2k downloads1mo agoHugging Face11ruggsea /infini-news-index INFINI-NEWS FM-Index 🔎 Live search API: these FM-indexes power a public search service — full-text search, n-gram counts, and document retrieval in the browser or via a keyless REST API, without building the index yourself — at infini-news.uni-graz.at (API reference). Pre-built FM-indexes (Burrows–Wheeler Transform + suffix array, built with infini-gram-mini, Liu et al. 2025) over the ruggsea/infini-news-corpus parquets. Enables exact, byte-level substring count and document… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-index.text-retrieval2 likes9k downloads5d agoHugging Face12ambrosfitz /19c_newspapers_images_altotabular100K<n<1M4 likes7.7k downloads3mo agoHugging Face13PleIAs /US-PD-Newspapers 🇺🇸 US Public Domain Newspapers 🇺🇸 US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library. With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining. Content As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.texttext-generation10M<n<100M50 likes7.5k downloads3y agoHugging Face14muse-bench /MUSE-News MUSE-News MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles… See the full description on the dataset page: https://huggingface.co/datasets/muse-bench/MUSE-News.text10K<n<100K4 likes6.3k downloads2y agoHugging Face15thegauravgiri /nepali-news-dataset 🇳🇵 Nepali News Dataset & NLP Corpus The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours. Repository: thegauravgiri/nepali-news-dataset Total Articles: 15,000+ full-text articles and growing Update Frequency: Every 4 hours via automated GitHub Actions pipelines Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding License: MIT License ⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.texttext-classification10K<n<100K1 likes6.1k downloads3h agoHugging Face16tmnam20 /Vietnamese-News Dataset Card for "VietnameseNewsparquet" More Information needed text1M<n<10M0 likes6k downloads3y agoHugging Face17SetFit /ag_newstext100K<n<1M10 likes5.1k downloads5y agoHugging Face18NealCaren /newspaper-pagesimage0 likes4.8k downloads2mo agoHugging Face19zeroshot /twitter-financial-news-sentiment Dataset Description The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment. The dataset holds 11,932 documents annotated with 3 labels: sentiments = { "LABEL_0": "Bearish", "LABEL_1": "Bullish", "LABEL_2": "Neutral" } The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.texttext-classification10K<n<100K179 likes4.3k downloads3y agoHugging Face20AlastairH /bbc-news-logger BBC News Surface Observations An independent longitudinal research dataset recording which stories appear on the BBC News front page and Most Read list, plus parsed snapshots of linked articles. This project is not affiliated with or endorsed by the BBC. Headlines, article text, and linked content remain subject to the BBC's terms and copyright. The collection is published for research, audit, and journalistic analysis; users are responsible for ensuring their use is lawful.… See the full description on the dataset page: https://huggingface.co/datasets/AlastairH/bbc-news-logger.time-series-forecasting3 likes4.2k downloads3d agoHugging Face21NealCaren /newspaper-ocr0 likes4.2k downloads4mo agoHugging Face22dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4k downloads1y agoHugging Face23vblagoje /cc_news Dataset Card for CC-News Dataset Summary CC-News dataset contains news articles from news sites all over the world. The data is available on AWS S3 in the Common Crawl bucket at /crawl-data/CC-NEWS/. This version of the dataset has been prepared using news-please - an integrated web crawler and information extractor for news.It contains 708241 English language news articles published between Jan 2017 and December 2019. It represents a small portion of the English… See the full description on the dataset page: https://huggingface.co/datasets/vblagoje/cc_news.imagetext-generation100K<n<1M70 likes3.6k downloads3y agoHugging Face24PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes3.5k downloads3y agoHugging Face25brian-learns /cdx-cc-news cdx-cc-news CDXj indexes for Common Crawl News Dataset. See also News Dataset Announcement I could not find indexes to Common Crawl News, so I decided to create some. Then I created a RocksDB Index and a lookup API. Direct url to API: https://brian-learns-cc-news-cdx-server.hf.space/lookup CLI for querying this index and retrieving archived web pages from CC-NEWS Common Crawl WARC files https://github.com/brian-learns/ccnget Directory Layout ├──… See the full description on the dataset page: https://huggingface.co/datasets/brian-learns/cdx-cc-news.tabular1B<n<10B8 likes3.4k downloads20d agoHugging Face26orionweller /cc_news_mds_incremental-tokens0 likes2.9k downloads2y agoHugging Face27hotchpotch /multilingual_cc_news hotchpotch/multilingual_cc_news Dataset Summary This dataset republishes multilingual CC-News data in a Hugging Face friendly layout with one subset per language. Source and transformation Original source datasets on the Hugging Face Hub: CloverSearch/cc-news-mutlilingual intfloat/multilingual_cc_news The intfloat version provides a loading script, but it can be difficult to use directly via the datasets library because it pulls raw JSONL files… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/multilingual_cc_news.text100M<n<1B0 likes2.9k downloads3mo agoHugging Face28idleengine /financial-news-multisource Multi-Source Financial & General News 🚀 57.1 MILLION ROWS OF NEWS CONTENT — one unified corpus for market-aware AI/ML I combined 24 public news datasets (many small on their own) into one consistent, ready-to-use layer so you don’t have to wrangle them yourself. Everything is normalized to a minimal schema (date, text, extra_fields) and shipped as Parquet shards per subset—streamable, DuckDB-friendly, and built with a trading date policy (this can be edited if folks see other… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/financial-news-multisource.texttext-classification10M<n<100M0 likes2.8k downloads1mo agoHugging Face29SetFit /bbc-news BBC News Topic Dataset Dataset on BBC News Topic Classification consisting of 2,225 articles published on the BBC News website corresponding during 2004-2005. Each article is labeled under one of 5 categories: business, entertainment, politics, sport or tech. Original source for this dataset: Derek Greene, Pádraig Cunningham, “Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering,” in Proc. 23rd International Conference on Machine learning (ICML’06)… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/bbc-news.texttext-classification1K<n<10K24 likes2.7k downloads2y agoHugging Face30nuuuwan /lk-news-chunkstabular100K<n<1M1 likes2.4k downloads2h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.