datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SP_500_Stocks_Data-ratios_news_price_10_yrsHi folks,
Here is a collection of data I have scraped or aggregated for most of the stocks in the S&P 500, including popular ones like Apple (AAPL).
It has the following data:
Daily news articles and sentiments on those articles collected over the last few years.
All quarterly stock fundamentals (ratios) for 10-20 years.
Stock price data (daily close) over the last 10-20 years.
Use it however you please for PERSONAL USAGE, but if you do leverage it to make some money; just remember me and… See the full description on the dataset page: https://huggingface.co/datasets/pmoe7/SP_500_Stocks_Data-ratios_news_price_10_yrs.crypto-news-coindesk-2020-2025
CoinDesk Cryptocurrency News Dataset (2020–2025)
This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics.
Time Period
January 1, 2020 – January 1, 2025
Content Overview
Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.fake-news-detection-dataset-EnglishThis is a cleaned and splitted version of this dataset (https://www.kaggle.com/datasets/sadikaljarif/fake-news-detection-dataset-english)
Labels:
Fake News: 0
Real News: 1
You can find the cleansing script at: https://github.com/ErfanMoosaviMonazzah/Fake-News-Detection
cs-230-news-v3ai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.finance-news-sentiment-35k
Finance News Sentiment 40k
39,965 English financial news headlines, collected from public Telegram
finance news-wire channels, labeled for 3-class sentiment (positive / negative / neutral) and a secondary
topic label, by two independent LLM judges from different model families with
an arbiter settling disputes.
A FinBERT model fine-tuned on this data reaches test accuracy 0.847 / macro F1
0.810: remehostingservices/finbert-finance-news-sentiment.
Code, training scripts and the… See the full description on the dataset page: https://huggingface.co/datasets/remehostingservices/finance-news-sentiment-35k.20-Newsgroups
20 Newsgroups
Train
Some measurable characteristics of the dataset:
D — number of documents
W — modality dictionary size (number of unique tokens)
len D — average document length in modality tokens (number of tokens)
len D uniq — average document length in unique modality tokens (number of unique tokens)
D
@lemmatized W
@lemmatized len D
@lemmatized len D uniq
@bigram W
@bigram len D
@bigram len D uniq
value
11301
1.0614e+06
93.9204
60.5687
213701
18.9099… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/20-Newsgroups.ViSL-News
ViSL-News
Dataset Summary
ViSL-News is a sentence-level Vietnamese Sign Language (VSL) dataset constructed from sign-interpreted Vietnamese news broadcasts.
The dataset was built from HTV Tin Tức videos published on YouTube during 2024–2025. Each sample consists of a sentence-level sign-language video clip paired with a Vietnamese text sentence.
ViSL-News was constructed using ViSL-Tool, a semi-automated framework designed for news videos that contain spoken… See the full description on the dataset page: https://huggingface.co/datasets/kha2612/ViSL-News.FinRL_BTC_news_signals
Overview
This news dataset is created for FinAI Contest 2025 Task 1 FinRL-DeepSeek for Crypto Trading. We collected BTC news for the training and testing period from different sources [1] [2]. For each news, we use the DeepSeek chat model to extract the sentiment score, risk level, and their correpsonding confidence level and one-sentence reasoning.
Column
Description
date_time
Timestamp of when the news article was published (in UTC).
title
Title of the news article.… See the full description on the dataset page: https://huggingface.co/datasets/SecureFinAI-Lab/FinRL_BTC_news_signals.SpatioTemporal-News-Corpus
Spatiotemporal News Dataset
Overview
This dataset contains approximately 1.2 million English-language news headlines and articles sourced from major outlets in the
United States, United Kingdom, Canada, and Australia.
Each entry is annotated with spatial (country of origin) and temporal (date of publication) contexts, designed for training spatiotemporal-aware sentence embeddings,
specifically our Space-Time-MiniLM-v0 model.
The dataset covers the period from January… See the full description on the dataset page: https://huggingface.co/datasets/Artur-B/SpatioTemporal-News-Corpus.turkish-disaster-news-geonlp
Turkish Disaster News GeoNLP Dataset
Dataset Summary
This dataset was created and submitted as part of the Uncharted Data Challenge by Adaption. The LLM-enhanced instruction pairs (turkish_earthquake_news.csv) were generated using Adaptive Data by Adaption — an AI-powered data adaptation platform.
The first open-source Turkish-language disaster news dataset with district-level geocoding, humanitarian category labels, and multi-dimensional damage classification.… See the full description on the dataset page: https://huggingface.co/datasets/FatmaElik/turkish-disaster-news-geonlp.viral_news_pairsThis dataset consists of popular news articles from google+, linkedin and facebook
In additin, a label column has been added to show the virality of the respectice title.
hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.economic-telegram-news-corpus-2025
Economic Telegram News Corpus 2025
A corpus of 31,292 Russian-language economic news posts collected from 7 major Telegram channels, spanning January 2024 to September 2025. The dataset supports research on economic narrative detection, topic classification, and information diffusion in social media.
Associated Paper
Going Viral: LLM-Based Modeling of Economic Narratives
Dataset Description
The raw collection contains 123,273 posts. The economic corpus was… See the full description on the dataset page: https://huggingface.co/datasets/bruhwalkk/economic-telegram-news-corpus-2025.Amharic-News-Text-classification-Dataset
An Amharic News Text classification Dataset
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.vn-provinces-newspaper-magazine-offices
Vietnam newspaper and magazine offices
Vietnam newspaper and magazine offices. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1071 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (102 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-newspaper-magazine-offices.mrd
NewSpace Market MRD
Typed, datasheet-sourced specifications for 2,486 spacecraft flight hardware products
across 24 categories — reaction wheels, star trackers, propulsion, radios, antennas,
structures, ground stations and more — with 19,011 values, each citing the vendor
datasheet page it came from.
The dataset exists because component selection is a task where a language model must not
improvise: a wheel that cannot take a 12 V bus is not a "close enough" answer. Three rules… See the full description on the dataset page: https://huggingface.co/datasets/NewSpaceMarket/mrd.Politifact_fake_newsrussian_oil_gas_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.french_financial_news
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/arcticgiant/french-financial-news
Context
This dataset contains around 41 500 french news from 11/2018 to 03/2021 scraped on a famous financial media website.
For ease of use I’v add English translation (Helsinki-NLP/opus-mt-fr-en) and sentiment analysis (VADER)
Analysis
The picture below show the effect of covid crisis on news sentiment (Purple) and CAC40 (Blue).
We see clearly a link between the news sentiment… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/french_financial_news.Indonesian-Health-Newsfake_news_combinedLabel Description
0 : Fake,
1 : Real
news_media_reliability
Reliability Estimation of News Media Sources: "Birds of a Feather Flock Together"
Dataset introduced in the paper "Reliability Estimation of News Media Sources: Birds of a Feather Flock Together" published in the NAACL 2024 main conference.
Similar to the news media bias and factual reporting dataset, this dataset consists of a collections of 5.33K new media domains names with reliability labels. Additionally, for some domains, there is also a human-provided reliability score… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_reliability.news-12factor
Dataset Card for news-12factor
Dataset Description
80+ news articles with url, title, body text, scored on 12 quality factors and assigned a single rank.
Languages
The text in the dataset is in English
Dataset Structure
[Needs More Information]
Source Data
URL data was scraped using news-please
Annotations
Articles were manually annotated by Alex on a 12-factor score card.
ro-offense-news
Dataset Card for "RO-News-Offense"
Dataset Summary
a novel Romanian language dataset for offensive message detection with manually
annotated comment from a local Romanian news website (stiri de cluj) into five classes:
non-offensive
targeted insults
racist
homophobic
sexist
Resulting in 4052 annotated messages
Languages
Romanian
Dataset Structure
Data Instances
An example of 'train' looks as follows.
{
'comment_id': 5… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/ro-offense-news.cs230-news-unfilteredPolitifact-fake-news-6-categories-for-llama3-1
Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News"
based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data
language:"
- en
license: llama3.1
global-newspaper-c9acb4
global-newspaper-c9acb4
Synthetic sensors test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/inouechiyo/global-newspaper-c9acb4.Sinhala-News-Source-classificationThis dataset contains Sinhala news headlines extracted from 9 news sources (websites) (Sri Lanka Army, Dinamina, GossipLanka, Hiru, ITN, Lankapuwath, NewsLK,
Newsfirst, World Socialist Web Site-Sinhala). This is a processed version of the corpus created by Sachintha, D., Piyarathna, L., Rajitha, C., and Ranathunga, S. (2021). Exploiting parallel corpora to improve multilingual embedding based document and sentence alignment. Single word sentences, invalid characters have been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Source-classification.Fake-News-ClassificationDevelop a machine learning program to identify when an article might be fake news. Run by the UTK Machine Learning Club.
This is the Dataset to the Fake-News-Classifier competition in Kaggle. There is a Test csv to check for predictions.
Citation
William Lifferth. (2018). Fake News. Kaggle. https://kaggle.com/competitions/fake-news
