CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shash42 /forecast-news Forecast News Deduplicated daily news corpus used by forecast-sim and future-sim. Snapshot 31,859,020 articles 3,463 daily partitions Coverage: 2016-08-26 through 2026-08-31 Snapshot published: 2026-09-18 Stored data size: approximately 158.5 GiB The Parquet files are the canonical complete representation. The repository also contains daily JSONL files where available and compact headline JSON files used by article-browsing workflows. Layout Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.text10M<n<100M2 likes36k downloads6d agoHugging Face02aisbergpublicorganization /telegram-news-ua-dataset Aisberg Telegram News UA A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.texttext-classification100K<n<1M4 likes19k downloads2h agoHugging Face03SetFit /20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs: The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date. We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.text10K<n<100K21 likes9.4k downloads5y agoHugging Face04SetFit /ag_newstext100K<n<1M10 likes5.1k downloads5y agoHugging Face05dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face06SetFit /bbc-news BBC News Topic Dataset Dataset on BBC News Topic Classification consisting of 2,225 articles published on the BBC News website corresponding during 2004-2005. Each article is labeled under one of 5 categories: business, entertainment, politics, sport or tech. Original source for this dataset: Derek Greene, Pádraig Cunningham, “Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering,” in Proc. 23rd International Conference on Machine learning (ICML’06)… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/bbc-news.texttext-classification1K<n<10K24 likes2.9k downloads2y agoHugging Face07BAAI /IndustryCorpus_news[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_news.texttext-generation100M<n<1B5 likes1.8k downloads1mo agoHugging Face08sh0416 /ag_newsAG's News Topic Classification Dataset Version 3, Updated 09/09/2015 ORIGIN AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search… See the full description on the dataset page: https://huggingface.co/datasets/sh0416/ag_news.texttext-classification100K<n<1M15 likes1.8k downloads4y agoHugging Face09heegyu /news-category-datasetDataset from https://www.kaggle.com/datasets/rmisra/news-category-dataset text100K<n<1M4 likes1.6k downloads4y agoHugging Face10titasmallick96 /daily-bio-newsaudion<1K0 likes1.5k downloads5h agoHugging Face11ZechengLi19 /CSL-News Summary This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News is a large-scale Chinese Sign Language dataset designed for developing robust sign language understanding models. Code: https://github.com/ZechengLi19/Uni-Sign Download Please refer to download script to download CSL_News. You can also download each file by wget, for instance: wget… See the full description on the dataset page: https://huggingface.co/datasets/ZechengLi19/CSL-News.textvideo-text-to-text100K<n<1M16 likes1.4k downloads1y agoHugging Face12Sachin21112004 /news-entertainment-datasettexttable-question-answeringn<1K6 likes934 downloads3h agoHugging Face13robotizac /News_Category_Dataset_v2text100K<n<1M1 likes789 downloads2y agoHugging Face14Sachin21112004 /news-politics-datasettext1K<n<10K0 likes698 downloads6h agoHugging Face15Sachin21112004 /news-education-datasettext10K<n<100K3 likes645 downloads3h agoHugging Face16Sachin21112004 /news-tech-datasettext10K<n<100K1 likes634 downloads3h agoHugging Face17Sachin21112004 /news-finance-datasettext10K<n<100K3 likes602 downloads3h agoHugging Face18jhu-clsp /news21-instructionstexttext-retrieval10K<n<100K1 likes522 downloads7mo agoHugging Face19KrossKinetic /SP500-Financial-News-Articles-Time-SeriesTextual Time Series Dataset for finetuning / pretraining. Json version of original dataset. Original Dataset : https://www.kaggle.com/datasets/skywalker290/financial-news-article-and-stock-trend-dataset?select=stock_data_articles.csv text1K<n<10K6 likes369 downloads2y agoHugging Face20common-pile /news News Description We scrape the news sites that publish content under CC BY or CC BY-SA according to opennewswire. These include 360info, Africa is a Country, Alt News, Balkan Diskurs, Factly, Freedom of the Press Foundation, Agenzia Fides, Global Voices, Meduza, Mekong Eye, Milwaukee Neighborhood News Service, Minority Africa, New Canadian Media, SciDev.Net, The Solutions Journalism Exchange, Tasnim News Agency… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/news.texttext-generation100K<n<1M2 likes317 downloads1y agoHugging Face21Superinn /CSL-News Summary This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News is a large-scale Chinese Sign Language dataset designed for developing robust sign language understanding models. Code: https://github.com/ZechengLi19/Uni-Sign Download Please refer to download script to download CSL_News. You can also download each file by wget, for instance: wget… See the full description on the dataset page: https://huggingface.co/datasets/Superinn/CSL-News.textvideo-text-to-text100K<n<1M0 likes309 downloads2mo agoHugging Face22davanstrien /newspaper_navigatorimageimage-to-text10M<n<100M0 likes304 downloads4y agoHugging Face23olm /gdelt-news-headlinestext10M<n<100M0 likes300 downloads4y agoHugging Face24Alwin-Yang /bmw-pressclub-news BMW PressClub News Dataset This dataset contains press releases and news articles scraped from BMW PressClub. Dataset Structure JSON Format (bmw_articles.json) { "scraped_at": "2025-12-17T10:00:00", "source": "https://www.press.bmwgroup.com/global/article", "count": 100, "articles": [ { "title": "BMW presents the new X5", "date": "17.12.2025", "article_type": "Press Release", "summary": "...", "tags": ["BMW X5", "SUV"]… See the full description on the dataset page: https://huggingface.co/datasets/Alwin-Yang/bmw-pressclub-news.texttext-generationn<1K0 likes275 downloads6mo agoHugging Face25raquiba /Sarcasm_News_HeadlinePast studies in Sarcasm Detection mostly make use of Twitter datasets collected using hashtag based supervision but such datasets are noisy in terms of labels and language. Furthermore, many tweets are replies to other tweets and detecting sarcasm in these requires the availability of contextual tweets. To overcome the limitations related to noise in Twitter datasets, this Headlines dataset for Sarcasm Detection is collected from two news website. TheOnion aims at producing sarcastic versions… See the full description on the dataset page: https://huggingface.co/datasets/raquiba/Sarcasm_News_Headline.text10K<n<100K6 likes248 downloads4y agoHugging Face26imthanhlv /binhvq_news21_rawtext10M<n<100M0 likes231 downloads5y agoHugging Face27ZechengLi19 /CSL-News_pose Summary This is the pose format dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News is a large-scale Chinese Sign Language dataset designed for developing robust sign language understanding models. Code: https://github.com/ZechengLi19/Uni-Sign Download Please refer to download script to download CSL_News. You can also download each file by wget, for instance: wget… See the full description on the dataset page: https://huggingface.co/datasets/ZechengLi19/CSL-News_pose.textvideo-text-to-text100K<n<1M2 likes202 downloads2y agoHugging Face28liuqi6777 /trec-newstexttext-retrieval100K<n<1M0 likes167 downloads1y agoHugging Face29ScoutieAutoML /russian-news-telegram-dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic News and Media, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the channel. views… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian-news-telegram-dataset.tabulartext-classification10K<n<100K6 likes164 downloads2y agoHugging Face30lincoln /newsquadfr Dataset Card for newsquadfr Dataset Summary newsquadfr is a small dataset created for Question Answering task. Contexts are paragraphs of articles extracted from nine online french newspaper during year 2020/2021. newsquadfr stands for Newspaper question answering dataset in french. inspired by Piaf and Squad dataset. 2 520 triplets context - question - answer. from datasets import load_dataset ds_name = 'lincoln/newsquadfr' # exemple 1 ds_newsquad =… See the full description on the dataset page: https://huggingface.co/datasets/lincoln/newsquadfr.tabularquestion-answering1K<n<10K2 likes151 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.