datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.earnings-call-transcriptslanguage:
en
tags:
finance
earnings-calls
transcripts
nlp
llm
rag
financial-analysis
license: other
pretty_name: Earnings Call Transcripts
size_categories:
- 10K<n<100K
Earnings Call Transcripts Dataset
A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages.
Dataset Overview
This dataset contains:
Company earnings call transcripts
Ticker symbols
Earnings quarters
Earnings years
Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.Stocks-Quarterly-Earnings
Stocks Quarterly Earnings
This dataset includes quarterly earnings report data for various stocks.
355,371 rows over 6,406 symbols, 8 columns, covering 1996-01-31 to 2026-07-31. Refreshed monthly.
Strategies Built on This Data
490 papers in the Papers With Backtest catalogue declare this dataset as an input. 456 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.20, and 32% clear a t-statistic of 1.96 on their own… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Quarterly-Earnings.sp500-earnings-transcripts
S&P 500 Earnings Call Transcripts
Dataset Description
This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals.
📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025.
Coverage Statistics
Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.earnings_callThe dataset reports a collection of earnings call transcripts, the related stock prices, and the sector index In terms of volume, there is a total of 188 transcripts, 11970 stock prices, and 1196 sector index values. Furthermore, all of these data originated in the period 2016-2020 and are related to the NASDAQ stock market. Furthermore, the data collection was made possible by Yahoo Finance and Thomson Reuters Eikon. Specifically, Yahoo Finance enabled the search for stock values and Thomson Reuters Eikon provided the earnings call transcripts. Lastly, the dataset can be used as a benchmark for the evaluation of several NLP techniques to understand their potential for financial applications. Moreover, it is also possible to expand the dataset by extending the period in which the data originated following a similar procedure.Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.earnings-call-data
S&P 500 earnings episodes (2005–2025)
Augmented release built on Bose345/sp500_earnings_transcripts (same transcript calendar span as that collection: 2005–2025). Static tabular data for supervised learning or RL-style experiments on earnings-call episodes. Each row is one company–quarter call, keyed by a stable episode_id, with long-form text (full earnings transcript, SEC press materials), pre-earnings price context, OHLCV anchors, SEC XBRL fundamentals (xbrl_* columns), and… See the full description on the dataset page: https://huggingface.co/datasets/RudrakshNanavaty/earnings-call-data.EarningsCall-Benchsp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.earnings-calendar
US Earnings Calendar
When a company announced its results, to the second, what those results
were, and which session could act on them.
454 611 earnings releases · 258 482 paired with reported figures ·
4 382 182 corporate events · 2003-04-25 to 2026-09-04
The pipeline lives in recipe/ at the same revision as the data.
See PIPELINE.md for the method.
This is not a calendar of future releases
Nothing here schedules an announcement. A row appears when EDGAR accepted… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/earnings-calendar.dia-earning21-all
Earnings 21
The Earnings 21 dataset ( also referred to as earnings21 ) is a 39-hour corpus of earnings calls containing entity dense speech from nine different financial sectors. This corpus is intended to benchmark automatic speech recognition (ASR) systems in the wild with special attention towards named entity recognition (NER).
This work has been recently accepted to Interspeech 2021!
File Format Overview
In the following section, we provide an overview of the file… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-earning21-all.stocks-earnings-income_statementpit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/jdecim/pit-earnings-call-qa.pit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/pit-earnings-call-qa.stk-earnings-transcripts2024-earnings-call-transcriptearnings_surprise
Earnings Surprise
Data Notice: This dataset provides academic research access with a 6-month data lag.
For real-time data access, please visit sov.ai to subscribe.
For market insights and additional subscription options, check out our newsletter at blog.sov.ai.
from datasets import load_dataset
df_earnings = load_dataset("sovai/earnings_surprise", split="train").to_pandas().set_index(["ticker","date"])
Data arrives late Friday night 11 pm - 12; the model also retrains weekly.… See the full description on the dataset page: https://huggingface.co/datasets/sovai/earnings_surprise.earnings-surprise-stocks
Dataset Card for "earnings-surprise-stocks"
More Information needed
earnings-call-data
S&P 500 earnings episodes (2005–2025)
Augmented release built on Bose345/sp500_earnings_transcripts (same transcript calendar span as that collection: 2005–2025). Static tabular data for supervised learning or RL-style experiments on earnings-call episodes. Each row is one company–quarter call, keyed by a stable episode_id, with long-form text (full earnings transcript, SEC press materials), pre-earnings price context, OHLCV anchors, SEC XBRL fundamentals (xbrl_* columns), and… See the full description on the dataset page: https://huggingface.co/datasets/Amaanaush/earnings-call-data.stocks-earnings-eps_estimateStocks-Weekly-EarningSurprise
Stocks Weekly Earnings Surprise
Weekly earnings surprise probabilities and outcomes for publicly traded companies.
2,204,032 rows over 6,067 symbols, 8 columns, covering 2016-12-30 to 2026-07-03. Refreshed monthly.
Why It Matters
This dataset supplies high-frequency earnings-surprise context for equity strategies by:
Pre-event positioning: Surprise probabilities guide sizing and hedging ahead of earnings announcements.
Post-event drift: Actual vs. estimated EPS… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Weekly-EarningSurprise.europe-ilo-ear-emtm-sex-cur-nb-median-monthly-earnings-of-employees-by-sex-and-cu
Median monthly earnings of employees by sex and currency | Europe (ILOSTAT)
🇪🇺 5,733 observations · 33 Europe countries · 1991–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 5,733 observations of Earnings data across 33 Europe countries, spanning 1991–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-ilo-ear-emtm-sex-cur-nb-median-monthly-earnings-of-employees-by-sex-and-cu.europe-ilo-ear-emta-sex-edu-nb-average-monthly-earnings-of-employees-by-sex-and-e
Average monthly earnings of employees by sex and education (local currency) | Europe (ILOSTAT)
🇪🇺 13,556 observations · 33 Europe countries · 1991–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 13,556 observations of Earnings data across 33 Europe countries, spanning 1991–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-ilo-ear-emta-sex-edu-nb-average-monthly-earnings-of-employees-by-sex-and-e.earnings-estimate-sp500earnings-sp500asia-ilo-ear-emtm-sex-cur-nb-median-monthly-earnings-of-employees-by-sex-and-cu
Median monthly earnings of employees by sex and currency | Asia (ILOSTAT)
🌏 2,388 observations · 27 Asia countries · 1996–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 2,388 observations of Earnings data across 27 Asia countries, spanning 1996–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ear-emtm-sex-cur-nb-median-monthly-earnings-of-employees-by-sex-and-cu.stocks-earnings-cash_flow_statementsp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.fairleap-driver-earnings-regression-500
Fairleap Driver Earnings Regression 500
📘 Dataset Overview
A small synthetic tabular dataset of ride-hailing driver work sessions, built for
the Fairleap AI project — a platform addressing income
uncertainty and wellbeing for Gojek/GOTO drivers in Indonesia. Each row is one work
session: when it happened, where, how long it ran, how many rides were completed, and
what it paid.
The dataset backs two regression targets: earnings (Indonesian Rupiah per session)
and… See the full description on the dataset page: https://huggingface.co/datasets/fairleap-ai/fairleap-driver-earnings-regression-500.
