CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zeroshot /twitter-financial-news-sentiment Dataset Description The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment. The dataset holds 11,932 documents annotated with 3 labels: sentiments = { "LABEL_0": "Bearish", "LABEL_1": "Bullish", "LABEL_2": "Neutral" } The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.texttext-classification10K<n<100K179 likes4.4k downloads3y agoHugging Face02zeroshot /twitter-financial-news-topic Dataset Description The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic. The dataset holds 21,107 documents annotated with 20 labels: topics = { "LABEL_0": "Analyst Update", "LABEL_1": "Fed | Central Banks", "LABEL_2": "Company | Product News", "LABEL_3": "Treasuries | Corporate Debt", "LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.texttext-classification10K<n<100K43 likes1.8k downloads3y agoHugging Face03shaurya03 /tech-news-dailytext10K<n<100K33 likes1.2k downloads1h agoHugging Face04Kushal0532 /news-political-bias-classification-datasetDataset actually from kaggle. Couldn't find it here so I uploaded it. text10K<n<100K0 likes888 downloads11mo agoHugging Face05madao33 /new-title-chinesetext1K<n<10K21 likes774 downloads4y agoHugging Face06daekeun-ml /naver-news-summarization-ko Naver-News-KO: A Korean News Summarization Dataset A Korean news summarization dataset of 27,400 (document, summary) pairs, crawled from Naver News over a ten-day window in July 2022. It was originally built for a Korean NLP hands-on lab and has been publicly hosted on the Hugging Face Hub since January 2023. A technical report documenting the collection protocol, corpus statistics, contamination analysis, and reproducible baselines is available on arXiv: arXiv:2607.20442.… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/naver-news-summarization-ko.textsummarization10K<n<100K65 likes619 downloads2mo agoHugging Face07batubayk /TR-News Citation If you use the dataset, please cite the paper: @article{10.1007/s10579-021-09568-y, year = {2022}, title = {{Abstractive text summarization and new large-scale datasets for agglutinative languages Turkish and Hungarian}}, author = {Baykara, Batuhan and Güngör, Tunga}, journal = {Language Resources and Evaluation}, issn = {1574-020X}, doi = {10.1007/s10579-021-09568-y}, pages = {1--35}} textsummarization100K<n<1M19 likes589 downloads4y agoHugging Face08galileo-ai /20_Newsgroups_Fixed Dataset Card for 20_Newsgroups_Fixed Dataset Summary This dataset is a version of the 20 Newsgroups dataset fixed with the help of the Galileo ML Data Intelligence Platform. In a matter of minutes, Galileo enabled us to uncover and fix a multitude of errors within the original dataset. In the end, we present this improved dataset as a new standard for natural language experimentation and benchmarking using the Newsgroups dataset. Curation Rationale This… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/20_Newsgroups_Fixed.texttext-classification10K<n<100K3 likes570 downloads4y agoHugging Face09genesisqu /fake-real-newstext10K<n<100K0 likes543 downloads4y agoHugging Face10pmoe7 /SP_500_Stocks_Data-ratios_news_price_10_yrsHi folks, Here is a collection of data I have scraped or aggregated for most of the stocks in the S&P 500, including popular ones like Apple (AAPL). It has the following data: Daily news articles and sentiments on those articles collected over the last few years. All quarterly stock fundamentals (ratios) for 10-20 years. Stock price data (daily close) over the last 10-20 years. Use it however you please for PERSONAL USAGE, but if you do leverage it to make some money; just remember me and… See the full description on the dataset page: https://huggingface.co/datasets/pmoe7/SP_500_Stocks_Data-ratios_news_price_10_yrs.tabular100K<n<1M30 likes542 downloads4y agoHugging Face11hanlincs /in1k_clip_qwen25vl_3b_224res_64tokens_new_pttabular1M<n<10M0 likes423 downloads1y agoHugging Face12hanlincs /in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_pttabular1M<n<10M0 likes393 downloads1y agoHugging Face13APProjects /new-york-layoffs-warn-act-notices-daily New York WARN Act layoff notices — every filing we hold since 2001, one CSV, rebuilt daily 6,515 New York WARN notices — every one this dataset holds, back to 2001 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-08-25 · state source last checked 2026-09-23T12:27Z · official source: New York Department of Labor — WARN notices. New York employers must file a WARN Act notice with the state before a qualifying mass layoff or plant… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/new-york-layoffs-warn-act-notices-daily.texttabular-classification1K<n<10K0 likes370 downloads2h agoHugging Face14APProjects /new-jersey-layoffs-warn-act-notices-daily New Jersey WARN Act layoff notices — every filing we hold since 2003, one CSV, rebuilt daily 2,330 New Jersey WARN notices — every one this dataset holds, back to 2003 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-01 · state source last checked 2026-09-23T12:23Z · official source: New Jersey Department of Labor and Workforce Development — WARN notices. New Jersey employers must file a WARN Act notice with the state before… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/new-jersey-layoffs-warn-act-notices-daily.tabulartabular-classification1K<n<10K0 likes360 downloads2h agoHugging Face15tum-nlp /neural-news-benchmark AI-generated News Detection Benchmark neural-news is a benchmark dataset designed for human/AI news authorship classification in English, Turkish, Hungarian, and Persian. Presented in Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian @ NLP for Positive Impact Workshop @ EMNLP2024. Dataset Details The dataset includes equal parts human-written and AI-generated news articles, raw and pre-processed. Curated… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/neural-news-benchmark.texttext-classification10K<n<100K4 likes344 downloads2y agoHugging Face16okite97 /news-data Dataset Card for news-data Dataset Summary The News Dataset is an English-language dataset containing just over 4k unique news articles scrapped from AriseTv- One of the most popular news television in Nigeria. Supported Tasks and Leaderboards It supports news article classification into different categories. Languages English Dataset Structure Data Instances ''' {'Title': 'Nigeria: APC Yet to Zone Party Positions Ahead of… See the full description on the dataset page: https://huggingface.co/datasets/okite97/news-data.texttext-classification1K<n<10K7 likes337 downloads4y agoHugging Face17gustavecortal /diverse_french_newstext100K<n<1M4 likes333 downloads3y agoHugging Face18maryamfakhari /crypto-news-coindesk-2020-2025 CoinDesk Cryptocurrency News Dataset (2020–2025) This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics. Time Period January 1, 2020 – January 1, 2025 Content Overview Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.imagetext-classification100K<n<1M2 likes324 downloads1mo agoHugging Face19mrm8488 /fake-newstext10K<n<100K1 likes312 downloads5y agoHugging Face20reilleo /new_deploy_big_actstextn<1K0 likes302 downloads11mo agoHugging Face21gopalkalpande /bbc-news-summary About Dataset Context Text summarization is a way to condense the large amount of information into a concise form by the process of selection of important information and discarding unimportant and redundant information. With the amount of textual information present in the world wide web the area of text summarization is becoming very important. The extractive summarization is the one where the exact sentences present in the document are used as summaries. The extractive… See the full description on the dataset page: https://huggingface.co/datasets/gopalkalpande/bbc-news-summary.text1K<n<10K23 likes293 downloads4y agoHugging Face22APProjects /new-mexico-layoffs-warn-act-notices-daily New Mexico WARN Act layoff notices — every filing we hold since 2016, one CSV, rebuilt daily 116 New Mexico WARN notices — every one this dataset holds, back to 2016 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-06-29 · state source last checked 2026-09-22T12:27Z. New Mexico employers must file a WARN Act notice with the state before a qualifying mass layoff or plant closing. This page is generated from those filings… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/new-mexico-layoffs-warn-act-notices-daily.tabulartabular-classificationn<1K0 likes290 downloads2h agoHugging Face23mariagrandury /fake_news_corpus_spanish Fake News Corpus Spanish Citation Gómez-Adorno, H., Posadas-Durán, J. P., Enguix, G. B., & Capetillo, C. P. (2021). Overview of FakeDeS at IberLEF 2021: Fake News Detection in Spanish Shared Task. Procesamiento del Lenguaje Natural, 67, 223-231. Aragón, M. E., Jarquín, H., Gómez, M. M. Y., Escalante, H. J., Villaseñor-Pineda, L., Gómez-Adorno, H., ... & Posadas-Durán, J. P. (2020, September). Overview of mex-a3t at iberlef 2020: Fake news and aggressiveness analysis in… See the full description on the dataset page: https://huggingface.co/datasets/mariagrandury/fake_news_corpus_spanish.texttext-classificationn<1K2 likes209 downloads2y agoHugging Face24reilleo /new_justify_eval_actstextn<1K0 likes205 downloads11mo agoHugging Face25ErfanMoosaviMonazzah /fake-news-detection-dataset-EnglishThis is a cleaned and splitted version of this dataset (https://www.kaggle.com/datasets/sadikaljarif/fake-news-detection-dataset-english) Labels: Fake News: 0 Real News: 1 You can find the cleansing script at: https://github.com/ErfanMoosaviMonazzah/Fake-News-Detection tabulartext-classification10K<n<100K5 likes201 downloads4y agoHugging Face26contemmcm /20_newsgroupstexttext-classification10K<n<100K0 likes195 downloads2y agoHugging Face27edaschau /bitcoin_newsBitcoin news scrapped from Yahoo Finance. Columns: time_unix the UNIX timestamp of the news (UTC) date_time UTC date and time text_matches the news articles are matched with keywords "BTC", "bitcoin", "crypto", "cryptocurrencies", "cryptocurrency". The list is the posititions the keywords appeared. title_matches keyword matches in title url the Yahoo Finance URL that the article from source the source if the news is cited from other source, not originally from Yahoo Finane source_url the outer… See the full description on the dataset page: https://huggingface.co/datasets/edaschau/bitcoin_news.textsummarization100K<n<1M17 likes195 downloads1y agoHugging Face28sergioburdisso /news_media_bias_and_factuality News Media Factual Reporting and Political Bias Dataset introduced in the paper "Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions" published in the CLEF 2024 main conference. Similar to the news media reliability dataset, this dataset consists of a collections of 4K new media domains names with political bias and factual reporting labels. Columns of the dataset: source: domain name bias: the political bias label. Values: "left"… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_bias_and_factuality.text1K<n<10K4 likes193 downloads2y agoHugging Face29SuryaKrishna02 /aya-telugu-news-articles Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.texttext-generation100K<n<1M6 likes185 downloads3y agoHugging Face30dsfsi /daily-news-dikgang Daily News Dikgang Give Feedback 📑: DSFSI Resource Feedback Form About dataset The dataset contains annotated categorised data from Dikgang - Daily News https://dailynews.gov.bw/news-list/srccategory/10. The data is in setswana. See the Data Statement for foll details. Disclaimer This dataset contains machine-readable data extracted from online news articles, from https://dailynews.gov.bw/news-list/srccategory/10, provided by the Botswana Government. While… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/daily-news-dikgang.texttext-classification1K<n<10K2 likes182 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.