datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.transcripts
My Weird Prompts — Transcript Corpus
Every published transcript from the My Weird Prompts
podcast, shaped for textual analysis: narrowed metadata, the full transcript, the
same transcript segmented into speaker turns, and per-episode text statistics.
5,318 episodes · 463,566 speaker turns. Rebuilt daily from the
production database.
Configs
from datasets import load_dataset
episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.sejm-committee-transcripts
Polish Sejm committee transcripts — full API coverage (terms 9 and 10)
Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm
of the Republic of Poland, parsed into individually attributed speaker turns.
Scope
Committees: all standing committees with zapis PDFs in the Sejm API.
Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17).
Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.measuring_cot_monitorability_transcripts
Measuring Chain-of-Thought Monitorability Transcripts
This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness.
We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.rlvr-reward-hacking-transcripts
RLVR reward-hacking full trajectories
This release contains 900 full held-out trajectories from three policies trained with
reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable
CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B
checkpoints. Each row preserves the task, tests, complete prompts, native
reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.altayduel-transcripts
🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2)
Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir.
📌 TL;DR
2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu.
439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.rlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.fomc-meeting-transcripts
FOMC Meeting Transcripts (1976–2020)
Full-text transcripts of 373 Federal Open Market Committee (FOMC) meetings, from March 1976 through December 2020, converted from the official PDF transcripts published by the Federal Reserve Board.
The FOMC is the body of the U.S. Federal Reserve System that sets monetary policy (the federal funds rate target, balance-sheet policy, etc.). Verbatim meeting transcripts are released to the public with a roughly five-year lag, which is why… See the full description on the dataset page: https://huggingface.co/datasets/brishen/fomc-meeting-transcripts.terence-mckenna-transcripts
Terence McKenna Transcripts
Catalog of machine-transcribed talks and interviews, mostly by Terence McKenna, built from 172 videos. Two tables:
talks — one row per video; full original transcript (including [SPEAKER_XX] diarization tags)
turns — one row per non-empty line; speaker tags stripped from text
Speaker labels are not consistent across files (no cross-video voice fingerprinting), so turns has no speaker column.
Load
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/paulalesius/terence-mckenna-transcripts.council-transcripts
Council Multi-Agent Deliberation Transcripts
Real deliberation transcripts from Council, a multi-agent orchestration skill for OpenClaw.
What This Is
Council routes tasks to specialized Grok persona agents — Workhorse (deep technical reasoning), Creative (novel ideas, chaos energy), and Speed (fast iteration) — then synthesizes their outputs into a final verdict via a Conductor. These transcripts capture the full deliberation process: prompts, per-persona responses, and… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/council-transcripts.podcast-transcripts
Podcast Transcripts Dataset
This dataset contains transcripts from Bitcoin and cryptocurrency podcasts,
processed by the belief-engines ETL pipeline.
Files
transcripts.parquet - Full episode transcripts with speaker diarization
transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings
Schema
transcripts.parquet
episode_id: Unique episode identifier
podcast_slug: Podcast name slug
episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.velvet-rope-playtest-transcripts
Velvet Rope Playtest Transcripts
Cleaned playtest transcripts for Velvet Rope, a Build Small Hackathon Gradio game where players talk past whimsical AI gatekeepers by reading moods and discovering each character's soft spot.
This dataset is published for the hackathon's sharing-is-caring badge. It contains 341 turn-level rows from 96 local playtest session files.
Files
data/playtest_transcripts.csv - table-friendly version.
data/playtest_transcripts.jsonl - one… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/velvet-rope-playtest-transcripts.frontier-ai-podcast-transcripts
Frontier AI Researcher Podcast Transcripts
Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode.
Contents
327 episodes
49,286 merged dialogue turns
5,595,982 English tokens using the o200k_base tokenizer
3,942,026 tokens in guest turns
Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.podcast-transcripts-cleaned-phase2
Podcast Transcripts Cleaned (Phase 2)
Updated 2026-07-18T23-19-55Z UTC.
Configs
Config
Rows
Description
episodes
4,304
Full-episode raw ASR → cleaned transcript (with episode_id / show_id)
chunks
5,646
Per-chunk raw → cleaned pairs
traces
287,271
Full LM traces (cleaner + evaluator): prompts, reasoning/CoT, tool outputs, pass/fail
Cleaned pairs (episodes / chunks)
Columns include episode_id, show_id, title, url, instruction… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/podcast-transcripts-cleaned-phase2.
