datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
earnings-call-transcriptslanguage:
en
tags:
finance
earnings-calls
transcripts
nlp
llm
rag
financial-analysis
license: other
pretty_name: Earnings Call Transcripts
size_categories:
- 10K<n<100K
Earnings Call Transcripts Dataset
A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages.
Dataset Overview
This dataset contains:
Company earnings call transcripts
Ticker symbols
Earnings quarters
Earnings years
Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.earnings_callThe dataset reports a collection of earnings call transcripts, the related stock prices, and the sector index In terms of volume, there is a total of 188 transcripts, 11970 stock prices, and 1196 sector index values. Furthermore, all of these data originated in the period 2016-2020 and are related to the NASDAQ stock market. Furthermore, the data collection was made possible by Yahoo Finance and Thomson Reuters Eikon. Specifically, Yahoo Finance enabled the search for stock values and Thomson Reuters Eikon provided the earnings call transcripts. Lastly, the dataset can be used as a benchmark for the evaluation of several NLP techniques to understand their potential for financial applications. Moreover, it is also possible to expand the dataset by extending the period in which the data originated following a similar procedure.earnings-call-data
S&P 500 earnings episodes (2005–2025)
Augmented release built on Bose345/sp500_earnings_transcripts (same transcript calendar span as that collection: 2005–2025). Static tabular data for supervised learning or RL-style experiments on earnings-call episodes. Each row is one company–quarter call, keyed by a stable episode_id, with long-form text (full earnings transcript, SEC press materials), pre-earnings price context, OHLCV anchors, SEC XBRL fundamentals (xbrl_* columns), and… See the full description on the dataset page: https://huggingface.co/datasets/RudrakshNanavaty/earnings-call-data.EarningsCall-Benchpit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/jdecim/pit-earnings-call-qa.pit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/pit-earnings-call-qa.2024-earnings-call-transcriptearnings-call-data
S&P 500 earnings episodes (2005–2025)
Augmented release built on Bose345/sp500_earnings_transcripts (same transcript calendar span as that collection: 2005–2025). Static tabular data for supervised learning or RL-style experiments on earnings-call episodes. Each row is one company–quarter call, keyed by a stable episode_id, with long-form text (full earnings transcript, SEC press materials), pre-earnings price context, OHLCV anchors, SEC XBRL fundamentals (xbrl_* columns), and… See the full description on the dataset page: https://huggingface.co/datasets/Amaanaush/earnings-call-data.earnings_call_transcript_autograder
Earnings Call LLM Insights
📚 Read the Full Story: For a deep dive into the methodology, the wildest moments we found, and key takeaways, check out the blog post:KnowTrend.ai: Auto-Grading Ten Years of Earnings Calls for Prescience and Delusion
This dataset contains LLM-generated analysis of ~70,000+ earnings call transcripts.
The analysis was performed using Kimi k2-0905-preview, focusing on extracting specific insights, prescient analyst questions, and management missteps.… See the full description on the dataset page: https://huggingface.co/datasets/knowtrendllc/earnings_call_transcript_autograder.
