datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.earnings-calls-qa
Lamini Earning Calls QA Dataset
Description
This dataset contains transcripts of earning calls for various companies, along with questions and answers related to the companies' financial performance and other relevant topics.
Format
The transcripts, questions, and answers are in the form of jsonlines files, with each json object in the file containing the transcript of an earning call for a single company.
Data Pipeline Code
The entire data pipeline… See the full description on the dataset page: https://huggingface.co/datasets/lamini/earnings-calls-qa.pit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/jdecim/pit-earnings-call-qa.pit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/pit-earnings-call-qa.2024-earnings-call-transcriptearnings_call_mono
Motley Fool Earnings Call Mono (Private)
This private dataset contains cleaned monolingual earnings-call text chunks prepared from:
Kaggle dataset: tpotterer/motley-fool-scraped-earnings-call-transcripts
Split
train: 135306 rows
Columns
id: chunk identifier
text: cleaned source text chunk
ticker: ticker symbol
exchange: exchange string from source metadata
date: call date string from source metadata
section: prepared or qa
Processing Summary… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/earnings_call_mono.earnings_10k
Dataset Summary
This dataset is curated to train (next-token) the LLM-ADE model(https://arxiv.org/abs/2404.13028), specifically designed to imbue it with financial domain expertise.
It consists of 75,849 sequences, amounting to approximately 16.8 million tokens, using the Llama tokenizer. We have deliberately unlabled the sequences wrt the company to reflect real world data and train the model to process knowledge from unlabelled data.
The data focuses on the 500 constituent… See the full description on the dataset page: https://huggingface.co/datasets/InvestmentResearchAI/earnings_10k.tigerbot-earning-pluginTigerbot 模型rethink时使用的外脑原始数据,财报类
共2500篇财报,抽取后按段落保存
发布时间区间为: 2022-02-28 至 2023-05-10
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-earning-plugin')
ai-earnings-2026earningsfinancial-earnings-dpoearning-dpo-sectioned
