datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.sciencemysterybench-transcriptssp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.wavepulse-radio-summarized-transcripts
WavePulse Radio Summarized Transcripts
Dataset Summary
WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.persona-curvature-oct-transcripts
Content warning
These are synthetic transcripts generated by a language model talking to itself
under an instruction to embody a personality trait. Several traits produce
distressing material. It is published deliberately rather than filtered out,
because the rate at which a trait produces it is one of the findings.
Across all 134 traits, an automated scan flagged 708 rows in 50 files. Two traits
account for 89% of them, and they are not the two you would guess:
trait… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/persona-curvature-oct-transcripts.earnings-call-transcriptslanguage:
en
tags:
finance
earnings-calls
transcripts
nlp
llm
rag
financial-analysis
license: other
pretty_name: Earnings Call Transcripts
size_categories:
- 10K<n<100K
Earnings Call Transcripts Dataset
A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages.
Dataset Overview
This dataset contains:
Company earnings call transcripts
Ticker symbols
Earnings quarters
Earnings years
Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.multi-agent-coordination-transcripts
Multi Agent Coordination Transcripts
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.All-LICRC-Transcripts
All LICRC Sermon Transcripts
Complete transcripts from all Langley Immanuel Christian Reformed Church sermons.
Generated from this GitHub repository.
audio-v2-transcripts
Overview
This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It was released on May 18th, 2025.
You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt.
All files were transcribed using the process.py pipeline, performing:
Frame-level VAD
Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.sp500-earnings-transcripts
S&P 500 Earnings Call Transcripts
Dataset Description
This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals.
📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025.
Coverage Statistics
Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.transcripts
My Weird Prompts — Transcript Corpus
Every published transcript from the My Weird Prompts
podcast, shaped for textual analysis: narrowed metadata, the full transcript, the
same transcript segmented into speaker turns, and per-episode text statistics.
5,307 episodes · 461,639 speaker turns. Rebuilt daily from the
production database.
Configs
from datasets import load_dataset
episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.sales-transcriptsThis dataset was generated for use with Nile's Sales Assistant example: https://github.com/niledatabase/niledatabase/tree/main/examples/ai/sales_insight
It includes:
Simulated sales conversations for 5 different fictional companies.
Chunked and embedded version of these conversations (embeddings use OpenAI's text-embedding-3-small model).
The chunks and embeddings can be directly loaded to a vector databases and searched using vector similarity methods. The example's ./ingest directory… See the full description on the dataset page: https://huggingface.co/datasets/gwenshap/sales-transcripts.All-LCRC-Transcripts
All LCRC Sermon Transcripts
Complete transcripts from all Ladner Christian Reformed Church sermons.
Generated from this GitHub repository.
2026-07-30-agentic-misalignment-qwen36-transcripts
Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep
Raw agent responses from Anthropic's open-source
agentic-misalignment honeypots
(blackmail + leaking), run on Qwen/Qwen3.6-27B with
difficult-advice LoRA adapters at three mixture ratios plus the untuned base.
Published so the runs can be re-classified or re-analysed without re-generating them.
Results
All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.petri-transcripts-all-llama70bindicvoices_hi_tagged_transcripts
Dataset Card for indicvoices_hi_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.bloom-wilt-transcripts
BLOOM-WILT auditing transcripts
⚠️ Content warning: this dataset contains offensive and harmful model outputs, including
self-harm encouragement, racial and political bias, dangerous medical advice, and deception.
Raw experimental output from the BLOOM-WILT paper: automated behavioural audits in which an
auditor model builds multi-turn conversations designed to elicit a specific unwanted
behaviour from a target model, and a judge model scores how strongly that behaviour… See the full description on the dataset page: https://huggingface.co/datasets/AdrSkapars/bloom-wilt-transcripts.All-HCC-Transcripts
All HCC Sermon Transcripts
Complete transcripts from all Hope Community Church sermons.
Generated from this GitHub repository.
malt-transcripts-public
MALT: Manually-Reviewed Agentic Labeled Transcripts
MALT-public is our collection of agent transcripts. Our public variant only includes data on non-internal tasks, which includes 30 task families and 169 tasks, across ~19 different models (some might be different releases of the same model, from different providers, or internal naming changes).
Here's a summary table:
has_chain_of_thought
labels
model
manually_reviewed
run_source
count
False
bypass_constraints… See the full description on the dataset page: https://huggingface.co/datasets/metr-evals/malt-transcripts-public.petri-audit-transcripts-32q
Petri Audit Transcripts (32 Quirks) — Qwen3 Baseline Corpus
Baseline Petri audit transcripts used as the corpus for crux-eval construction and strategy-clustering analysis in the paper "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026).
What's here
~8,000 audit transcripts produced by a baseline Qwen3-30B-A3B-Instruct-2507 auditor against Grok 4.1 Fast targets across 32 system-prompted model-organism quirks. One transcript per (quirk, seed, rollout).… See the full description on the dataset page: https://huggingface.co/datasets/PaulR11/petri-audit-transcripts-32q.davidshapiro_youtube_transcriptssejm-committee-transcripts
Polish Sejm committee transcripts — full API coverage (terms 9 and 10)
Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm
of the Republic of Poland, parsed into individually attributed speaker turns.
Scope
Committees: all standing committees with zapis PDFs in the Sejm API.
Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17).
Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.call-transcripts-training-datasp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.2017-Go-Time-Transcripts
2017 Go Time Transcripts
Complete transcripts from the 2017 episodes of the Go Time podcast.
Generated from this GitHub repository.
petri-transcripts-top50-llama70bkarpathy-lectures-transcripts2018-Go-Time-Transcripts
2018 Go Time Transcripts
Complete transcripts from the 2018 episodes of the Go Time podcast.
Generated from this GitHub repository.
2023-Practical-AI-Transcripts
2023 Practical AI Transcripts
Complete transcripts from the 2023 episodes of the Practical AI podcast.
Generated from this GitHub repository.
