CoolFace
Datasetpublic

cheruvo/financial-news-sentiment

Cheruvo Financial News Sentiment 25,548 financial news headlines from around the world, matched to 51 stocks and cryptocurrencies and scored from −1 to +1 from an investor's point of view. Collected 13 May – 19 September 2026. Data from The GDELT Project. Sentiment scores computed by Cheruvo. Why this exists I published an article claiming that news sentiment does not predict price, it follows it. That is an assertion you can either believe or check, and until now… See the full description on the dataset page: https://huggingface.co/datasets/cheruvo/financial-news-sentiment.

sourceHugging Faceotherupdated 5d agoView on Hugging Face
1likes37downloads
Dataset Card

Cheruvo Financial News Sentiment

25,548 financial news headlines from around the world, matched to 51 stocks and cryptocurrencies and scored from −1 to +1 from an investor's point of view. Collected 13 May – 19 September 2026.

Data from [The GDELT Project](https://www.gdeltproject.org/). Sentiment scores computed by [Cheruvo](https://cheruvo.com).

Why this exists

I published an article claiming that news sentiment does not predict price, it follows it. That is an assertion you can either believe or check, and until now you could only believe it. This is the file that lets you check.

What the dataset does not show

Read this before the columns, because it is the part that gets skipped.

Using the stable-period rows of this dataset, on Bitcoin, over 44 usable days and 5,285 stories: daily sentiment does not predict the next return at any horizon tested. At seven days the Spearman rank correlation is −0.001. Tested with a permutation test and a block bootstrap over block lengths 2, 3, 5, 7 and 10, reporting the worst block length, with the significance threshold fixed at 0.017 (0.05 across three horizons) before looking at any result.

What does hold is the opposite direction: today's sentiment tracks the price move of the previous two to seven days, rho ≈ 0.61, and that survives the block bootstrap. On Ethereum the same test fails, but with 926 stories against Bitcoin's 5,285 that is an underpowered replication rather than a failed one.

So if you are here to build a trading signal, this dataset is evidence against the thing you are about to try, from the person who collected it. You are very welcome to show me I got it wrong. That is the point of publishing it.

Columns

columnmeaning
tickerthe symbol the story was matched to
data_pubblicazionepublication timestamp, ISO 8601 with timezone
fonteGDELT · <domain>
linguasource language as declared by GDELT (three-letter code), when present
titolothe headline, as the outlet wrote it
urllink to the original article
sentiment−1 to +1, four decimals
origine_punteggiollm (Groq/Llama), vader, or av
periodo_raccoltawhich collection regime the row belongs to — see below

A story that names several tickers appears once per ticker.

The three collection regimes, and why you must not ignore them

The rows are not homogeneous, and the differences are ours, not the market's.

regimedatesrowsrows/day
pre-riforma13 May – 6 Aug3,409~40
filtro-asimmetrico7 Aug – 15 Aug5,048~561
stabile16 Aug – 19 Sep17,091~488

Density. There is a fourteenfold jump on 7 August. That is the day real collection started; everything before it is a thin historical backfill. If you plot story volume over the whole file you will see a wall in early August and it means nothing about the world.

Bias. Until 16 August the relevance filter recognised profits in five languages and had no word for losses in any of them. Measured at the time, the stories it discarded averaged −0.095 against +0.084 for the ones it kept. The sentiment series before that date is shifted upward by construction.

I kept those rows rather than deleting them quietly, because a reader who cannot see what is missing cannot judge it. If you are measuring anything about sentiment levels or their relationship to prices, use periodo_raccolta == "stabile" and nothing else. Any result that spans 7 or 16 August measures my bug fixes.

Coverage, including where it is bad

Ten best-covered tickers: BTC-USD (5,705), NVDA (2,577), AAPL (1,382), GOOGL (1,230), ETH-USD (1,180), XRP-USD (1,150), META (1,053), AMZN (994), MSFT (917), JPM (786).

Seventeen tickers have fewer than 100 rows across four months. A daily mean over a handful of articles is an anecdote, and they are thin for two different reasons that should not be confused.

Not collected. NFLX (7 rows) was never in the collection vocabulary. Its handful of rows are rebound mentions: stories gathered for other tickers that happen to name Netflix. It is not a statement about how much Netflix is in the news.

Collected, but the name is an ordinary word. The rest were searched for and came back thin, and for many of them the reason is the search term itself:

tickertermrows
APT-USDAptos11
NEAR-USDNEAR11
ARB-USDArbitrum14
MATIC-USDPolygon18
STMMI.MISTMicroelectronics23
LTC-USDLitecoin24
ATOM-USDCosmos27
DOT-USDPolkadot27
AVAX-USDAvalanche33
UNI-USDUniswap35
SHIB-USDShiba48
XLM-USDStellar64
LINK-USDChainlink68
SAN.MCSantander80
OP-USDOptimism89
AIR.PAAirbus92

NEAR, Cosmos, Avalanche, Stellar, Optimism, Polygon and Arbitrum are English words before they are assets. A headline reading "Fashion Styles Spur Optimism" is not about a layer-2 network, and the relevance filter correctly refuses it. So for those coins the honest sentence is not "there is little news about this asset". It is "we cannot reliably tell news about this asset apart from news that uses this word", which is a different and more useful warning.

The European listings (STMMI.MI, SAN.MC, AIR.PA) are thin for a third reason: GDELT does not index enough European financial press for them.

How the scores were produced

Headlines are scored by a large language model (Llama via Groq) prompted to judge the story from an investor's point of view, with VADER as a fallback when the model is unavailable. origine_punteggio records which produced each row.

The prompt was chosen by scoring 50 hand-labelled headlines against the candidates rather than by reading the outputs and picking the nicer one.

Scores are opinions of a model, not ground truth. Syndicated rewrites of the same story are deduplicated in the analysis but are present here as separate rows, since which outlets picked a story up is itself information.

Licence and attribution

The underlying news metadata comes from The GDELT Project, whose terms permit unlimited use, including commercial, and explicitly permit redistribution:

You may redistribute, rehost, republish, and mirror any of the GDELT datasets in any form. However, any use or redistribution of the data must include a citation to the GDELT Project and a link to this website.

That condition is binding on you too, so if you redistribute this dataset or anything derived from it, carry the attribution line at the top of this card.

Only GDELT-sourced rows are included. Cheruvo also holds rows from the ECB, ESMA, SEC EDGAR and Alpha Vantage; those are deliberately excluded, because their licences differ. The ECB and ESMA require modifications to be declared, and computing a sentiment score is a modification. Alpha Vantage granted written permission for Cheruvo's own use, which is not a redistribution licence. Bundling three licences under one attribution line would force you to trust rather than check.

Citation

bibtex
@misc{cheruvo2026sentiment,
  title  = {Cheruvo Financial News Sentiment},
  author = {Santacaterina, Michele},
  year   = {2026},
  note   = {Data from The GDELT Project (https://www.gdeltproject.org/)},
  url    = {https://cheruvo.com}
}

The code that produced this file is backend/esporta_dataset.py in the Cheruvo repository, and the test that decides whether a result may be reported is backend/verifica_segnale.py.