datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.pit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/jdecim/pit-earnings-call-qa.pit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/pit-earnings-call-qa.2024-earnings-call-transcriptai-earnings-2026financial-earnings-dpoearning-dpo-sectioned
