CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fancyzhx /ag_news Dataset Card for "ag_news" Dataset Summary AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc)… See the full description on the dataset page: https://huggingface.co/datasets/fancyzhx/ag_news.texttext-classification100K<n<1M196 likes90k downloads3y agoHugging Face02open-index /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.texttext-generation10M<n<100M343 likes48k downloads1mo agoHugging Face03ruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B39 likes46k downloads7d agoHugging Face04shash42 /forecast-news Forecast News Deduplicated daily news corpus used by forecast-sim and future-sim. Snapshot 31,859,020 articles 3,463 daily partitions Coverage: 2016-08-26 through 2026-08-31 Snapshot published: 2026-09-18 Stored data size: approximately 158.5 GiB The Parquet files are the canonical complete representation. The repository also contains daily JSONL files where available and compact headline JSON files used by article-browsing workflows. Layout Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.text10M<n<100M2 likes35k downloads5d agoHugging Face05aisbergpublicorganization /telegram-news-ua-dataset Aisberg Telegram News UA A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.texttext-classification100K<n<1M4 likes17k downloads18m agoHugging Face06AlphaDojo /dojo_stock_news Languages: 简体中文 · English dojo_stock_news — Stock News Overview Financial news linked to individual stocks: headline, summary, source, publish time, and URL. Files File Description data.parquet Full news archive Key Fields Field Description symbol Associated stock symbol (primary query key) title Headline description Summary body publish_date Publish date (YYYY-MM-DD or locale-specific text)… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_stock_news.text1M<n<10M0 likes14k downloads16m agoHugging Face07SetFit /20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs: The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date. We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.text10K<n<100K21 likes9.4k downloads5y agoHugging Face08institutional /institutional-newspapers-bplgated 📰 Institutional Newspapers: Boston Public Library A structured dataset derived from the Boston Public Library's public domain newspapers collection, produced by the Institutional Data Initiative in collaboration with Boston Public Library. 1,473,635 public domain newspaper scans, published between 1795 and 1930 83,147,041 individual crops segmented from those scans 16.3 billion o200k_base tokens of VLM OCR text, and 14.7 billion from Tesseract Data for each crop: bbox… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-newspapers-bpl.image1M<n<10M10 likes9.2k downloads1mo agoHugging Face09ambrosfitz /19c_newspapers_images_altotabular100K<n<1M4 likes7.8k downloads3mo agoHugging Face10PleIAs /US-PD-Newspapers 🇺🇸 US Public Domain Newspapers 🇺🇸 US-PD-Newspapers is an agregation of all the archives of US newspapers digitized by the Library of Congress for the Chronicling America digital library. With nearly 100 billion words, it is one of the largest open corpus in the United States. All the materials are now part of the public domain and have no intellectual property rights remaining. Content As of January 2024, the collection contains nearly 21 millions unique newspaper… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/US-PD-Newspapers.texttext-generation10M<n<100M50 likes7.2k downloads3y agoHugging Face11thegauravgiri /nepali-news-dataset 🇳🇵 Nepali News Dataset & NLP Corpus The comprehensive, open-access Nepali & English News Dataset for NLP and Machine Learning, automatically aggregated and updated every 4 hours. Repository: thegauravgiri/nepali-news-dataset Total Articles: 15,000+ full-text articles and growing Update Frequency: Every 4 hours via automated GitHub Actions pipelines Languages: Nepali (np / ne) and English (en) in clean UTF-8 Devanagari encoding License: MIT License ⚡ Free… See the full description on the dataset page: https://huggingface.co/datasets/thegauravgiri/nepali-news-dataset.texttext-classification10K<n<100K1 likes6.6k downloads5h agoHugging Face12PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes6.3k downloads3y agoHugging Face13muse-bench /MUSE-News MUSE-News MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles… See the full description on the dataset page: https://huggingface.co/datasets/muse-bench/MUSE-News.text10K<n<100K4 likes6.1k downloads2y agoHugging Face14tmnam20 /Vietnamese-News Dataset Card for "VietnameseNewsparquet" More Information needed text1M<n<10M0 likes6k downloads3y agoHugging Face15SetFit /ag_newstext100K<n<1M10 likes5.1k downloads5y agoHugging Face16zeroshot /twitter-financial-news-sentiment Dataset Description The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment. The dataset holds 11,932 documents annotated with 3 labels: sentiments = { "LABEL_0": "Bearish", "LABEL_1": "Bullish", "LABEL_2": "Neutral" } The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.texttext-classification10K<n<100K179 likes4.5k downloads3y agoHugging Face17dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face18brian-learns /cdx-cc-news cdx-cc-news CDXj indexes for Common Crawl News Dataset. See also News Dataset Announcement I could not find indexes to Common Crawl News, so I decided to create some. Then I created a RocksDB Index and a lookup API. Direct url to API: https://brian-learns-cc-news-cdx-server.hf.space/lookup CLI for querying this index and retrieving archived web pages from CC-NEWS Common Crawl WARC files https://github.com/brian-learns/ccnget Directory Layout ├──… See the full description on the dataset page: https://huggingface.co/datasets/brian-learns/cdx-cc-news.tabular1B<n<10B8 likes3.6k downloads22d agoHugging Face19vblagoje /cc_news Dataset Card for CC-News Dataset Summary CC-News dataset contains news articles from news sites all over the world. The data is available on AWS S3 in the Common Crawl bucket at /crawl-data/CC-NEWS/. This version of the dataset has been prepared using news-please - an integrated web crawler and information extractor for news.It contains 708241 English language news articles published between Jan 2017 and December 2019. It represents a small portion of the English… See the full description on the dataset page: https://huggingface.co/datasets/vblagoje/cc_news.imagetext-generation100K<n<1M70 likes3.4k downloads3y agoHugging Face20hotchpotch /multilingual_cc_news hotchpotch/multilingual_cc_news Dataset Summary This dataset republishes multilingual CC-News data in a Hugging Face friendly layout with one subset per language. Source and transformation Original source datasets on the Hugging Face Hub: CloverSearch/cc-news-mutlilingual intfloat/multilingual_cc_news The intfloat version provides a loading script, but it can be difficult to use directly via the datasets library because it pulls raw JSONL files… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/multilingual_cc_news.text100M<n<1B0 likes2.9k downloads3mo agoHugging Face21SetFit /bbc-news BBC News Topic Dataset Dataset on BBC News Topic Classification consisting of 2,225 articles published on the BBC News website corresponding during 2004-2005. Each article is labeled under one of 5 categories: business, entertainment, politics, sport or tech. Original source for this dataset: Derek Greene, Pádraig Cunningham, “Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering,” in Proc. 23rd International Conference on Machine learning (ICML’06)… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/bbc-news.texttext-classification1K<n<10K24 likes2.9k downloads2y agoHugging Face22idleengine /financial-news-multisource Multi-Source Financial & General News 🚀 57.1 MILLION ROWS OF NEWS CONTENT — one unified corpus for market-aware AI/ML I combined 24 public news datasets (many small on their own) into one consistent, ready-to-use layer so you don’t have to wrangle them yourself. Everything is normalized to a minimal schema (date, text, extra_fields) and shipped as Parquet shards per subset—streamable, DuckDB-friendly, and built with a trading date policy (this can be edited if folks see other… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/financial-news-multisource.texttext-classification10M<n<100M0 likes2.8k downloads1mo agoHugging Face23nuuuwan /lk-news-chunkstabular100K<n<1M1 likes2.5k downloads1h agoHugging Face24GonzaloA /fake_news TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: https://huggingface.co/spaces/huggingface/datasets-tagging annotations_creators: - no-annotation language_creators: - found language: - en license: - unknown multilinguality: - monolingual size_categories: - 30k<n<50k source_datasets: - original task_categories: - text-classification task_ids: - fact-checking - intent-classification pretty_name: GonzaloA / Fake News Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/GonzaloA/fake_news.tabular10K<n<100K28 likes2.2k downloads4y agoHugging Face25JulesBelveze /tldr_news Dataset Card for tldr_news Dataset Summary The tldr_news dataset was constructed by collecting daily tech newsletters from TLDR. For every piece of news, the title, content, category, section, and source URLs were extracted. The dataset now covers multiple TLDR newsletters including AI, tech, crypto, and other categories. This dataset can be used for various NLP tasks including: Headline generation Text summarization News categorization Content classification by… See the full description on the dataset page: https://huggingface.co/datasets/JulesBelveze/tldr_news.textsummarization10K<n<100K26 likes2.1k downloads9mo agoHugging Face26Helsinki-NLP /news_commentary Dataset Card for OPUS News-Commentary Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/news_commentary.texttranslation1M<n<10M39 likes2k downloads3y agoHugging Face27RealTimeData /bbc_news_alltime RealTimeData Monthly Collection - BBC News This datasets contains all news articles from BBC News that were created every months from 2017 to current. To access articles in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/bbc_news_alltime', '2020-02') This will give you all BBC news articles that were created in 2020-02. Want to crawl the data by your own? Please head to LatestEval for the crawler scripts. Credit… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_news_alltime.image100K<n<1M51 likes2k downloads1y agoHugging Face28oitnews /newsdataiotext10K<n<100K0 likes1.8k downloads18h agoHugging Face29zeroshot /twitter-financial-news-topic Dataset Description The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic. The dataset holds 21,107 documents annotated with 20 labels: topics = { "LABEL_0": "Analyst Update", "LABEL_1": "Fed | Central Banks", "LABEL_2": "Company | Product News", "LABEL_3": "Treasuries | Corporate Debt", "LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.texttext-classification10K<n<100K43 likes1.8k downloads3y agoHugging Face30ctoraman /sozcu-news-2014(This dataset contains raw text, which are unlabeled.) 1,656 Turkish news articles from Sözcü Newspaper (http://www.sozcu.com.tr) between December 20, 2013, and March 11, 2014. GitHub Repo: https://github.com/BilkentInformationRetrievalGroup/TUBITAK113E249/ If you would like to use any material in this repository, please cite this paper: Toraman, C. and Can, F. (2017), Discovering story chains: A framework based on zigzagged search and news actors. Journal of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/sozcu-news-2014.texttext-generation1K<n<10K0 likes1.8k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.