datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.earnings-call-transcriptslanguage:
en
tags:
finance
earnings-calls
transcripts
nlp
llm
rag
financial-analysis
license: other
pretty_name: Earnings Call Transcripts
size_categories:
- 10K<n<100K
Earnings Call Transcripts Dataset
A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages.
Dataset Overview
This dataset contains:
Company earnings call transcripts
Ticker symbols
Earnings quarters
Earnings years
Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.sp500-earnings-transcripts
S&P 500 Earnings Call Transcripts
Dataset Description
This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals.
📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025.
Coverage Statistics
Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.transcripts
My Weird Prompts — Transcript Corpus
Every published transcript from the My Weird Prompts
podcast, shaped for textual analysis: narrowed metadata, the full transcript, the
same transcript segmented into speaker turns, and per-episode text statistics.
5,339 episodes · 466,298 speaker turns. Rebuilt daily from the
production database.
Configs
from datasets import load_dataset
episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.sejm-committee-transcripts
Polish Sejm committee transcripts — full API coverage (terms 9 and 10)
Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm
of the Republic of Poland, parsed into individually attributed speaker turns.
Scope
Committees: all standing committees with zapis PDFs in the Sejm API.
Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17).
Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.measuring_cot_monitorability_transcripts
Measuring Chain-of-Thought Monitorability Transcripts
This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness.
We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.stk-earnings-transcriptsvaani-all-eng-transcripts-concatcollab-arena-v0-transcripts
Collaboration Arena v0 — transcripts & annotations
⚠️ This dataset is LIVE and GROWING until the experiment completes. Configs and counts update as Qwen cells and the 32B tier land; see the commit history for the update log.
ℹ️ E5 frontier model = claude-opus-4-8 (E1-E4 frontier = claude-fable-5). Fable safety-refused ~50% of E5 turns (stop_reason=refusal), concentrated on the honest cross-check seats, so E5's frontier arm was run on Opus instead; team and solo are BOTH Opus… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/collab-arena-v0-transcripts.swe_chat_transcriptsfomc-meeting-transcripts
FOMC Meeting Transcripts (1976–2020)
Full-text transcripts of 373 Federal Open Market Committee (FOMC) meetings, from March 1976 through December 2020, converted from the official PDF transcripts published by the Federal Reserve Board.
The FOMC is the body of the U.S. Federal Reserve System that sets monetary policy (the federal funds rate target, balance-sheet policy, etc.). Verbatim meeting transcripts are released to the public with a roughly five-year lag, which is why… See the full description on the dataset page: https://huggingface.co/datasets/brishen/fomc-meeting-transcripts.Kisan_Call_Centre_Transcriptshome-telecom-callcenter-transcripts-viewer
Home/Telecom Call Center Transcript Viewer V2
This public dataset is a flattened viewer copy of one source archive from AIxBlock/92k-real-world-call-center-scripts-english:
home_ervice_inbound&telecom _outbound.zip
It was converted so Hugging Face Data Studio can display the contents as regular Parquet tables.
Splits
default/train: one row per transcript, all 3,239 transcripts from the source zip.
turns/turns: one row per inferred timestamp/sentence turn, linked by… See the full description on the dataset page: https://huggingface.co/datasets/yevgeniy03/home-telecom-callcenter-transcripts-viewer.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.terence-mckenna-transcripts
Terence McKenna Transcripts
Catalog of machine-transcribed talks and interviews, mostly by Terence McKenna, built from 172 videos. Two tables:
talks — one row per video; full original transcript (including [SPEAKER_XX] diarization tags)
turns — one row per non-empty line; speaker tags stripped from text
Speaker labels are not consistent across files (no cross-video voice fingerprinting), so turns has no speaker column.
Load
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/paulalesius/terence-mckenna-transcripts.seamless-interact-canary-transcripts
Seamless Interact - Canary Transcripts
Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries.
Dataset Description
This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.urfunnyv2_transcriptscbai_deployment_transcriptsafrica-synth-retail-and-ecommerce-call-center-transcripts-nigeria
Call Center Transcripts | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-call-center-transcripts-nigeria.toronto-city-council-transcriptsKisan_Call_Centre_Transcriptsindic-short-utterance-transcripts
Indic Short-Utterance Transcripts (1–3 words), 10 languages
39,140 transcripts of 1–3 word utterances across 10 Indic languages, filtered from
ai4bharat/Rasa. Text only — no audio.
Built to answer a specific question: which short phrases exist in open Indic speech corpora,
and how many are in the register a voice assistant actually speaks? TTS models trained on
sentence-level read speech degrade on very short inputs because such corpora contain almost no
short utterances. This… See the full description on the dataset page: https://huggingface.co/datasets/OmS1ngh/indic-short-utterance-transcripts.garchen_rinpoche_inference_transcriptsindic-voices-scribe-v2-transcriptspodcast-transcripts
Podcast Transcripts Dataset
This dataset contains transcripts from Bitcoin and cryptocurrency podcasts,
processed by the belief-engines ETL pipeline.
Files
transcripts.parquet - Full episode transcripts with speaker diarization
transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings
Schema
transcripts.parquet
episode_id: Unique episode identifier
podcast_slug: Podcast name slug
episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.Podcast-Transcripts-Deduped
Podcast Transcripts (Deduped)
Phase 1 deduplicated subset of hudsongouge/Podcast-Transcripts-Raw.
Updated 2026-07-19T06-28-19Z UTC.
Dedup summary
Metric
Value
Raw input rows
102,374
Keepers
99,035
Discarded
3,339
Exact discarded
882
MinHash discarded
2457
Jaccard threshold
0.88
Two-pass Phase 1:
Exact match on canonical audio source key (YouTube ID / Omny clip / audio filename)
MinHash + LSH on normalized transcript text (Jaccard ≈… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Deduped.fed_transcripts_cleaned_slim
