datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment.twitter-financial-news-topic
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic.
The dataset holds 21,107 documents annotated with 20 labels:
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.financial-news-headlines
Financial News Headlines Dataset
A synthetic dataset of 10,038 financial news headlines with sentiment, sector, and topic labels — designed for NLP tasks like sentiment analysis, text classification, and semantic search.
📊 Dataset Overview
This dataset was generated using google/flan-t5-base from HuggingFace for a Data Science course project. Each headline is paired with rich metadata including sector, sentiment, topic, and company name.
Features
Column… See the full description on the dataset page: https://huggingface.co/datasets/KalsusEvening/financial-news-headlines.Indian_Financial_News
Dataset Card for Dataset Name
The FinancialNewsSentiment_26000 dataset comprises 26,000 rows of financial news articles related to the Indian market. It features four columns: URL, Content (scrapped content), Summary (generated using the T5-base model), and Sentiment Analysis (gathered using the GPT add-on for Google Sheets). The dataset is designed for sentiment analysis tasks, providing a comprehensive view of sentiments expressed in financial news.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/kdave/Indian_Financial_News.financial-news-sentimentData source: CSMAR
Link for raw data: https://www.heywhale.com/mw/dataset/5e577a780e2b66002c2561a9/content
Data Description:
2329 news titles with annotated labels (0:Negative, 1:Neutral, 2:Positive)
financial_newsfrench_financial_news
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/arcticgiant/french-financial-news
Context
This dataset contains around 41 500 french news from 11/2018 to 03/2021 scraped on a famous financial media website.
For ease of use I’v add English translation (Helsinki-NLP/opus-mt-fr-en) and sentiment analysis (VADER)
Analysis
The picture below show the effect of covid crisis on news sentiment (Purple) and CAC40 (Blue).
We see clearly a link between the news sentiment… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/french_financial_news.twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class… See the full description on the dataset page: https://huggingface.co/datasets/ROLEX-2007/twitter-financial-news-sentiment.Financial_News_Translation_Spanish_Finetune
Overview of the Financial News Translation Dataset for OpenAI Model Fine-tuning
Introduction:
This dataset has been curated with the primary objective of fine-tuning varioyus language models to effectively translate financial news content embedded in HTML format. The intention is to enhance the language model's proficiency in accurately and contextually translating financial information for a global audience in a production envionrment.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Benzinga/Financial_News_Translation_Spanish_Finetune.tradingview_msn_financial_news_1kfinancial-news-sentiment
Cheruvo Financial News Sentiment
25,548 financial news headlines from around the world, matched to 51 stocks and
cryptocurrencies and scored from −1 to +1 from an investor's point of view.
Collected 13 May – 19 September 2026.
Data from The GDELT Project. Sentiment
scores computed by Cheruvo.
Why this exists
I published an article claiming that news sentiment does not predict price, it
follows it. That is an assertion you can either believe or check, and until now… See the full description on the dataset page: https://huggingface.co/datasets/cheruvo/financial-news-sentiment.financial-news-by-ticker
Financial News by Tickers
This is the dataset for the financial news parsed from three russian news outlets: RIA Novosti, RBC, Kommersant.ru. The news were parsed for 20 tickers as follows:
LKOH, SBER, ROSN, GAZP, VTBR, YDEX, PLZL, T, NVTK, X5, GMKN, MGNT, ALRS, AFLT, CHMF, NLMK, MOEX, SNGSP, MTSS, PIKK
Overall, the dataset contains 5031 entries, the detailed stats by ticker and source:
ticker
source
count
till_dt
from_dt
AFLT
kommersant
100
2026-05-29 12:04:11.000… See the full description on the dataset page: https://huggingface.co/datasets/complicat9d/financial-news-by-ticker.twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class… See the full description on the dataset page: https://huggingface.co/datasets/KaiyuanLai/twitter-financial-news-sentiment.sentiment_analysis_financial_news_datacombined_financial_phrasebank_twitter_news_sentimentfinancial_newsmarket-insight-benchmarks
Financial News Market Insight Bridge Benchmarks
Benchmark dataset of 20 financial news article cases with individual scores for AI visibility, content discovery, topic matching, search visibility, article organization, and finance topic mapping.
Built by FinancialNews.it.com.
Dataset Description
This dataset contains benchmark data for a content visibility and AI discovery bot helping Financial News articles gain greater discoverability across AI platforms… See the full description on the dataset page: https://huggingface.co/datasets/financialnews/market-insight-benchmarks.FinancialNewsThis dataset contains the Fictional Financial news based on RavenPack.
Two jsonl files (df_neg.jsonl & df_pos.jsonl) can be directly used for OpenAI / OpenPipe finetuning.
(This is what you need for replication)
More detailed data at df_pos_generated and df_neg_generated file. The real_firm files are out-of-sample cues.
Us_Financial_news_dataset
US Financial News Dataset (Cleaned)
This dataset contains cleaned and structured financial news article data collected between 2013 and 2018. It includes only the article title, text, published date, language, and selected metadata like named entity counts — with all source website identifiers, URLs, and engagement metadata removed.
✅ What's Included
title: News headline
text: Full article body text
published: ISO-formatted date of publication
❌ What's Not… See the full description on the dataset page: https://huggingface.co/datasets/gowthamgoli/Us_Financial_news_dataset.bist-dp-lstm-trading-turkish_financial_news
turkish_financial_news
Turkish financial news corpus with sentiment labels
Dataset Details
Format: json
Size: ~50MB compressed
Language: Turkish (labels), Numeric (data)
License: MIT
Created: 2025-08-27
Dataset Structure
BIST Historical Data
Symbols: BIST 30 index stocks
Timeframes: 1m, 5m, 15m, 60m, 1d
Features: OHLCV + 131 technical indicators
Date Range: 2019-2024
Technical Indicators
Trend: SMA, EMA, MACD, Bollinger Bands… See the full description on the dataset page: https://huggingface.co/datasets/rsmctn/bist-dp-lstm-trading-turkish_financial_news.Indian_Financial_News
Dataset Card for Dataset Name
The FinancialNewsSentiment_26000 dataset comprises 26,000 rows of financial news articles related to the Indian market. It features four columns: URL, Content (scrapped content), Summary (generated using the T5-base model), and Sentiment Analysis (gathered using the GPT add-on for Google Sheets). The dataset is designed for sentiment analysis tasks, providing a comprehensive view of sentiments expressed in financial news.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/pranali96/Indian_Financial_News.twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/arabianpost/twitter-financial-news-sentiment.twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/ShuweiHou/twitter-financial-news-sentiment.twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class… See the full description on the dataset page: https://huggingface.co/datasets/SS171/twitter-financial-news-sentiment.financialNewsindian_financial_news_42kINDIAN FINANCIAL NEWS DATASET (42K)
42,214 Indian financial news headlines with sentiment labels and taxonomy-enriched annotations.
CONTENTS
Total rows: 42,214
Event-labeled rows: 2,979 (with event_id, macro_signal, sector impacts)
Sentiment-only rows: 39,235 (labeled by AION-Sentiment-IN-v3)
Synthetic data: 200 rows for macro_inr_appreciation (rupee appreciation with negative sentiment)
COLUMNS
headline: News headline text
event_id: Taxonomy event identifier (95 events, empty for… See the full description on the dataset page: https://huggingface.co/datasets/pranali96/indian_financial_news_42k.investing_financial_news_headlines
Dataset Sources
Investing.com
General_Financial_News_AltaredAltared version of : https://huggingface.co/datasets/lukecarlate/general_financial_news
Removed None data type
license: cc0-1.0
task_categories:
- text-classification
language:
- en
tags:
- finance
pretty_name: General Financial News Altared
size_categories:
- 10K<n<100K
twitter-financial-news-sentiment
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their sentiment.
The dataset holds 11,932 documents annotated with 3 labels:
sentiments = {
"LABEL_0": "Bearish",
"LABEL_1": "Bullish",
"LABEL_2": "Neutral"
}
The data was collected using the Twitter API. The current dataset supports the multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/abhivsep1/twitter-financial-news-sentiment.financial_news
