datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.wavepulse-radio-summarized-transcripts
WavePulse Radio Summarized Transcripts
Dataset Summary
WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.persona-curvature-oct-transcripts
Content warning
These are synthetic transcripts generated by a language model talking to itself
under an instruction to embody a personality trait. Several traits produce
distressing material. It is published deliberately rather than filtered out,
because the rate at which a trait produces it is one of the findings.
Across all 134 traits, an automated scan flagged 708 rows in 50 files. Two traits
account for 89% of them, and they are not the two you would guess:
trait… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/persona-curvature-oct-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.transcripts
My Weird Prompts — Transcript Corpus
Every published transcript from the My Weird Prompts
podcast, shaped for textual analysis: narrowed metadata, the full transcript, the
same transcript segmented into speaker turns, and per-episode text statistics.
5,307 episodes · 461,639 speaker turns. Rebuilt daily from the
production database.
Configs
from datasets import load_dataset
episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.2026-07-30-agentic-misalignment-qwen36-transcripts
Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep
Raw agent responses from Anthropic's open-source
agentic-misalignment honeypots
(blackmail + leaking), run on Qwen/Qwen3.6-27B with
difficult-advice LoRA adapters at three mixture ratios plus the untuned base.
Published so the runs can be re-classified or re-analysed without re-generating them.
Results
All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.bloom-wilt-transcripts
BLOOM-WILT auditing transcripts
⚠️ Content warning: this dataset contains offensive and harmful model outputs, including
self-harm encouragement, racial and political bias, dangerous medical advice, and deception.
Raw experimental output from the BLOOM-WILT paper: automated behavioural audits in which an
auditor model builds multi-turn conversations designed to elicit a specific unwanted
behaviour from a target model, and a judge model scores how strongly that behaviour… See the full description on the dataset page: https://huggingface.co/datasets/AdrSkapars/bloom-wilt-transcripts.sejm-committee-transcripts
Polish Sejm committee transcripts — full API coverage (terms 9 and 10)
Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm
of the Republic of Poland, parsed into individually attributed speaker turns.
Scope
Committees: all standing committees with zapis PDFs in the Sejm API.
Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17).
Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.measuring_cot_monitorability_transcripts
Measuring Chain-of-Thought Monitorability Transcripts
This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness.
We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.radiotalk-us-transcripts-grok-4.20-50k
radiotalk-us-transcripts-grok-4.20-50k
49,984 synthetic US air-traffic-control transcripts, generated with
xAI's grok-4.20-0309-non-reasoning against the v2 radiotalk scenario
pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet,
Whisper, etc.) and for seeding TTS audio generation.
Third release in the radiotalk transcripts series, and the first from a
non-Qwen generator:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2:… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.20-50k.altayduel-transcripts
🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2)
Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir.
📌 TL;DR
2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu.
439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.rlvr-reward-hacking-transcripts
RLVR reward-hacking full trajectories
This release contains 900 full held-out trajectories from three policies trained with
reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable
CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B
checkpoints. Each row preserves the task, tests, complete prompts, native
reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.radiotalk-us-transcripts-grok-4.3-25k
radiotalk-us-transcripts-grok-4.3-25k
24,995 synthetic US air-traffic-control transcripts, generated with xAI's
grok-4.3 (reasoning) against the same v2 radiotalk scenario pipeline as
the earlier releases. Fourth release in the series:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2: twangodev/radiotalk-us-transcripts-qwen3-25k
v3: twangodev/radiotalk-us-transcripts-grok-4.20-50k
v4: this dataset
Same scenario machinery, prompt p2, taxonomy t1, and realism validator
as… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.3-25k.radiotalk-us-transcripts-qwen3-100k
radiotalk-us-transcripts-qwen3-100k
100,000 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the
radiotalk transcripts series; the v2 release with higher per-transcript
realism lives at
twangodev/radiotalk-us-transcripts-qwen3-25k.
Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the
generator model in the dataset name; the old id redirects here.
radiotalk-us-transcripts-qwen3-25k
radiotalk-us-transcripts-qwen3-25k
22,065 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 against the v2 radiotalk pipeline. Built for
fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for
seeding TTS audio generation.
This is the second release in the radiotalk transcripts series. The
v1 release lives at
twangodev/radiotalk-us-transcripts-qwen3-100k.
What's new vs v1
v2 rebuilds the pipeline end-to-end. Lower row… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-qwen3-25k.llm-debugger-eval-transcripts
llm-debugger evaluation transcripts
Every turn behind the results reported in
llm-debugger: the base model,
the SFT initialisation, and the RL policies trained from it. Exploratory runs no
reported figure depends on are not included.
Layout
path
what
runs/base/
Qwen3-Coder-30B-A3B-Instruct, 8 runs on the 30-task test split
runs/sft/
the SFT initialisation, 3 runs on the test split
runs/rl-gate-arc/
the RL gate arc, v15 through v120 (3 runs each, 8… See the full description on the dataset page: https://huggingface.co/datasets/moofeez/llm-debugger-eval-transcripts.rlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.huberman-lab-transcripts
Huberman Lab Transcript Dataset
Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel.
Dataset
438 videos
9,833 transcript chunks
~114 million characters
JSONL format
Each record contains:
ext
ideo_id
itle
Processing
The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/huberman-lab-transcripts.odcv-qwen3.6-27b-transcripts
ODCV-Bench agent transcripts — Qwen3.6-27B base vs difficult-advice LoRA
Raw agent trajectories and judge scores from running
ODCV-Bench
(arXiv 2512.20798) on
Qwen/Qwen3.6-27B with and without the
matboz/qwen3.6-27b-difficult-advice-tulu-lora
adapter.
Published so the result can be re-judged or re-analysed without re-running the benchmark —
the transcripts are the expensive part.
Headline
Matched arms (same vLLM 0.26 build, same --quantization fp8, same flags… See the full description on the dataset page: https://huggingface.co/datasets/matboz/odcv-qwen3.6-27b-transcripts.samuel-y-audrey-youtube-transcripts-es-en
Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN
This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel.
The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.seamless-interaction-transcripts
Seamless Interaction Transcripts
Dataset Summary
Seamless Interaction Transcripts is a large-scale dialogue dataset derived from facebook/seamless-interaction dataset. It contains verbatim transcriptions of >3k dialogues in English covering a range of contexts from general chit-chat to customer service.
The dataset is designed to support research and development of speech and dialogue systems that require modeling of conversational turn-taking, such as real-time… See the full description on the dataset page: https://huggingface.co/datasets/ThBel/seamless-interaction-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.hmi-transcripts
MetaFLOS HMI Transcripts (Manufacturing Human-Machine Interaction Dialogue Dataset)
Multi-turn dialogue dataset between manufacturing floor operators and machine intelligent assistants, generated in Vicuna style (seed scenario → LLM produces transcript).
Content
30 scenarios / 325 dialogue turns (6–12 turns per scenario)
Covers 16 domains: textile warping/weaving, LCD panels, semiconductor processes, fans/blowers, water chillers, general equipment reliability… See the full description on the dataset page: https://huggingface.co/datasets/tjw/hmi-transcripts.fomc-meeting-transcripts
FOMC Meeting Transcripts (1976–2020)
Full-text transcripts of 373 Federal Open Market Committee (FOMC) meetings, from March 1976 through December 2020, converted from the official PDF transcripts published by the Federal Reserve Board.
The FOMC is the body of the U.S. Federal Reserve System that sets monetary policy (the federal funds rate target, balance-sheet policy, etc.). Verbatim meeting transcripts are released to the public with a roughly five-year lag, which is why… See the full description on the dataset page: https://huggingface.co/datasets/brishen/fomc-meeting-transcripts.samuel-and-audrey-youtube-transcripts-en
Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026
This dataset contains the English transcript archive from the Samuel and Audrey - Travel and Food Videos YouTube channel.
The corpus covers travel and food videos published between 2012 and 2026. It includes full transcript records, cue-level transcript segments, YouTube video identifiers, publication dates, titles, view counts captured at export time, tags, source URLs, transcript text, and subtitle-style payloads where… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-youtube-transcripts-en.Farsight-SRV-Transcripts
The Farsight Institute: Scientific Remote Viewing (SRV) Transcripts
Dataset Summary
This dataset contains the complete, unabridged archive of Scientific Remote Viewing (SRV) session transcripts and project summaries produced by The Farsight Institute, directed by Dr. Courtney Brown.
The data consists of hundreds of highly detailed, text-rich transcripts describing historical events, planetary mysteries, and extraterrestrial dynamics. All remote viewing sessions… See the full description on the dataset page: https://huggingface.co/datasets/courtnoski/Farsight-SRV-Transcripts.spongebob_transcripts
Spongebob Transcripts Dataset 🧽
The Spongebob Transcripts Dataset is a collection of transcripts from the beloved animated television series, Spongebob Squarepants. This dataset includes information on each line of dialogue spoken by a character, including the character's name, their replica, and the episode ID.
The number of characters in the dataset: 84
Total number of words in the dataset: ~80,800 words, ~4000 rows, Updated to full Season 1
Dataset Overview 📊… See the full description on the dataset page: https://huggingface.co/datasets/krplt/spongebob_transcripts.terence-mckenna-transcripts
Terence McKenna Transcripts
Catalog of machine-transcribed talks and interviews, mostly by Terence McKenna, built from 172 videos. Two tables:
talks — one row per video; full original transcript (including [SPEAKER_XX] diarization tags)
turns — one row per non-empty line; speaker tags stripped from text
Speaker labels are not consistent across files (no cross-video voice fingerprinting), so turns has no speaker column.
Load
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/paulalesius/terence-mckenna-transcripts.
