datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.twitter-financial-news-topic
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic.
The dataset holds 21,107 documents annotated with 20 labels:
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.tech-news-dailynews-political-bias-classification-datasetDataset actually from kaggle.
Couldn't find it here so I uploaded it.
TR-News
Citation
If you use the dataset, please cite the paper:
@article{10.1007/s10579-021-09568-y,
year = {2022},
title = {{Abstractive text summarization and new large-scale datasets for agglutinative languages Turkish and Hungarian}},
author = {Baykara, Batuhan and Güngör, Tunga},
journal = {Language Resources and Evaluation},
issn = {1574-020X},
doi = {10.1007/s10579-021-09568-y},
pages = {1--35}}
fake-real-newsnaver-news-summarization-ko
Naver-News-KO: A Korean News Summarization Dataset
A Korean news summarization dataset of 27,400 (document, summary) pairs, crawled from
Naver News over a ten-day window in July 2022. It was originally built for a
Korean NLP hands-on lab and has been publicly hosted on the Hugging Face Hub since January 2023.
A technical report documenting the collection protocol, corpus statistics, contamination analysis, and
reproducible baselines is available on arXiv: arXiv:2607.20442.… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/naver-news-summarization-ko.20_Newsgroups_Fixed
Dataset Card for 20_Newsgroups_Fixed
Dataset Summary
This dataset is a version of the 20 Newsgroups dataset fixed with the help of the Galileo ML Data Intelligence Platform. In a matter of minutes, Galileo enabled us to uncover and fix a multitude of errors within the original dataset. In the end, we present this improved dataset as a new standard for natural language experimentation and benchmarking using the Newsgroups dataset.
Curation Rationale
This… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/20_Newsgroups_Fixed.SP_500_Stocks_Data-ratios_news_price_10_yrsHi folks,
Here is a collection of data I have scraped or aggregated for most of the stocks in the S&P 500, including popular ones like Apple (AAPL).
It has the following data:
Daily news articles and sentiments on those articles collected over the last few years.
All quarterly stock fundamentals (ratios) for 10-20 years.
Stock price data (daily close) over the last 10-20 years.
Use it however you please for PERSONAL USAGE, but if you do leverage it to make some money; just remember me and… See the full description on the dataset page: https://huggingface.co/datasets/pmoe7/SP_500_Stocks_Data-ratios_news_price_10_yrs.crypto-news-coindesk-2020-2025
CoinDesk Cryptocurrency News Dataset (2020–2025)
This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics.
Time Period
January 1, 2020 – January 1, 2025
Content Overview
Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.fake-newsnews-data
Dataset Card for news-data
Dataset Summary
The News Dataset is an English-language dataset containing just over 4k unique news articles scrapped from AriseTv- One of the most popular news television in Nigeria.
Supported Tasks and Leaderboards
It supports news article classification into different categories.
Languages
English
Dataset Structure
Data Instances
'''
{'Title': 'Nigeria: APC Yet to Zone Party Positions Ahead of… See the full description on the dataset page: https://huggingface.co/datasets/okite97/news-data.diverse_french_newsneural-news-benchmark
AI-generated News Detection Benchmark
neural-news is a benchmark dataset designed for human/AI news authorship classification in English, Turkish, Hungarian, and Persian.
Presented in Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian @ NLP for Positive Impact Workshop @ EMNLP2024.
Dataset Details
The dataset includes equal parts human-written and AI-generated news articles, raw and pre-processed.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/neural-news-benchmark.bbc-news-summary
About Dataset
Context
Text summarization is a way to condense the large amount of information into a concise form by the process of selection of important information and discarding unimportant and redundant information. With the amount of textual information present in the world wide web the area of text summarization is becoming very important. The extractive summarization is the one where the exact sentences present in the document are used as summaries. The extractive… See the full description on the dataset page: https://huggingface.co/datasets/gopalkalpande/bbc-news-summary.fake_news_corpus_spanish
Fake News Corpus Spanish
Citation
Gómez-Adorno, H., Posadas-Durán, J. P., Enguix, G. B., & Capetillo, C. P. (2021). Overview of FakeDeS at IberLEF 2021: Fake News Detection in Spanish Shared Task. Procesamiento del Lenguaje Natural, 67, 223-231.
Aragón, M. E., Jarquín, H., Gómez, M. M. Y., Escalante, H. J., Villaseñor-Pineda, L., Gómez-Adorno, H., ... & Posadas-Durán, J. P. (2020, September). Overview of mex-a3t at iberlef 2020: Fake news and aggressiveness analysis in… See the full description on the dataset page: https://huggingface.co/datasets/mariagrandury/fake_news_corpus_spanish.fake-news-detection-dataset-EnglishThis is a cleaned and splitted version of this dataset (https://www.kaggle.com/datasets/sadikaljarif/fake-news-detection-dataset-english)
Labels:
Fake News: 0
Real News: 1
You can find the cleansing script at: https://github.com/ErfanMoosaviMonazzah/Fake-News-Detection
bitcoin_newsBitcoin news scrapped from Yahoo Finance.
Columns:
time_unix the UNIX timestamp of the news (UTC)
date_time UTC date and time
text_matches the news articles are matched with keywords "BTC", "bitcoin", "crypto", "cryptocurrencies", "cryptocurrency". The list is the posititions the keywords appeared.
title_matches keyword matches in title
url the Yahoo Finance URL that the article from
source the source if the news is cited from other source, not originally from Yahoo Finane
source_url the outer… See the full description on the dataset page: https://huggingface.co/datasets/edaschau/bitcoin_news.news_media_bias_and_factuality
News Media Factual Reporting and Political Bias
Dataset introduced in the paper "Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions" published in the CLEF 2024 main conference.
Similar to the news media reliability dataset, this dataset consists of a collections of 4K new media domains names with political bias and factual reporting labels.
Columns of the dataset:
source: domain name
bias: the political bias label. Values: "left"… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_bias_and_factuality.20_newsgroupsdaily-news-dikgang
Daily News Dikgang
Give Feedback 📑: DSFSI Resource Feedback Form
About dataset
The dataset contains annotated categorised data from Dikgang - Daily News https://dailynews.gov.bw/news-list/srccategory/10. The data is in setswana.
See the Data Statement for foll details.
Disclaimer
This dataset contains machine-readable data extracted from online news articles, from https://dailynews.gov.bw/news-list/srccategory/10, provided by the Botswana Government. While… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/daily-news-dikgang.aya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.fake_or_real_news
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Trinisha/fake_or_real_news.cs-230-news-v3covid_fake_newsConstraint@AAAI2021 - COVID19 Fake News Detection in English
@misc{patwa2020fighting,
title={Fighting an Infodemic: COVID-19 Fake News Dataset},
author={Parth Patwa and Shivam Sharma and Srinivas PYKL and Vineeth Guptha and Gitanjali Kumari and Md Shad Akhtar and Asif Ekbal and Amitava Das and Tanmoy Chakraborty},
year={2020},
eprint={2011.03327},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
ai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.finance-news-sentiment-35k
Finance News Sentiment 40k
39,965 English financial news headlines, collected from public Telegram
finance news-wire channels, labeled for 3-class sentiment (positive / negative / neutral) and a secondary
topic label, by two independent LLM judges from different model families with
an arbiter settling disputes.
A FinBERT model fine-tuned on this data reaches test accuracy 0.847 / macro F1
0.810: remehostingservices/finbert-finance-news-sentiment.
Code, training scripts and the… See the full description on the dataset page: https://huggingface.co/datasets/remehostingservices/finance-news-sentiment-35k.spanish-fake-news-fixed
Spanish Fake News Fixed
Este dataset contiene noticias etiquetadas en español, reparado para corregir saltos de línea internos.
Nigeria.NewsBBC_NEWS
