datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.wavepulse-radio-summarized-transcripts
WavePulse Radio Summarized Transcripts
Dataset Summary
WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.persona-curvature-oct-transcripts
Content warning
These are synthetic transcripts generated by a language model talking to itself
under an instruction to embody a personality trait. Several traits produce
distressing material. It is published deliberately rather than filtered out,
because the rate at which a trait produces it is one of the findings.
Across all 134 traits, an automated scan flagged 708 rows in 50 files. Two traits
account for 89% of them, and they are not the two you would guess:
trait… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/persona-curvature-oct-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.transcripts
My Weird Prompts — Transcript Corpus
Every published transcript from the My Weird Prompts
podcast, shaped for textual analysis: narrowed metadata, the full transcript, the
same transcript segmented into speaker turns, and per-episode text statistics.
5,307 episodes · 461,639 speaker turns. Rebuilt daily from the
production database.
Configs
from datasets import load_dataset
episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.2026-07-30-agentic-misalignment-qwen36-transcripts
Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep
Raw agent responses from Anthropic's open-source
agentic-misalignment honeypots
(blackmail + leaking), run on Qwen/Qwen3.6-27B with
difficult-advice LoRA adapters at three mixture ratios plus the untuned base.
Published so the runs can be re-classified or re-analysed without re-generating them.
Results
All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.Deltarune-Complete-Transcript-Cleaned
Deltarune Chapters 1–4 Dataset
Fan-made transcript dataset covering Deltarune Chapters 1 through 4. Processed from video playthroughs and cross-referenced with game data. Intended to provide LLMs with structured narrative context for a game whose content is underrepresented in training corpora.
Why This Exists
As of early 2026, major LLMs (including models with training cutoffs past July 2025) fail to recall basic plot details of Deltarune Chapters 3 and 4 despite their… See the full description on the dataset page: https://huggingface.co/datasets/Deltarunefan/Deltarune-Complete-Transcript-Cleaned.bloom-wilt-transcripts
BLOOM-WILT auditing transcripts
⚠️ Content warning: this dataset contains offensive and harmful model outputs, including
self-harm encouragement, racial and political bias, dangerous medical advice, and deception.
Raw experimental output from the BLOOM-WILT paper: automated behavioural audits in which an
auditor model builds multi-turn conversations designed to elicit a specific unwanted
behaviour from a target model, and a judge model scores how strongly that behaviour… See the full description on the dataset page: https://huggingface.co/datasets/AdrSkapars/bloom-wilt-transcripts.sejm-committee-transcripts
Polish Sejm committee transcripts — full API coverage (terms 9 and 10)
Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm
of the Republic of Poland, parsed into individually attributed speaker turns.
Scope
Committees: all standing committees with zapis PDFs in the Sejm API.
Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17).
Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.measuring_cot_monitorability_transcripts
Measuring Chain-of-Thought Monitorability Transcripts
This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness.
We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.radiotalk-us-transcripts-grok-4.20-50k
radiotalk-us-transcripts-grok-4.20-50k
49,984 synthetic US air-traffic-control transcripts, generated with
xAI's grok-4.20-0309-non-reasoning against the v2 radiotalk scenario
pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet,
Whisper, etc.) and for seeding TTS audio generation.
Third release in the radiotalk transcripts series, and the first from a
non-Qwen generator:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2:… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.20-50k.altayduel-transcripts
🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2)
Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir.
📌 TL;DR
2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu.
439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.rlvr-reward-hacking-transcripts
RLVR reward-hacking full trajectories
This release contains 900 full held-out trajectories from three policies trained with
reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable
CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B
checkpoints. Each row preserves the task, tests, complete prompts, native
reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.radiotalk-us-transcripts-grok-4.3-25k
radiotalk-us-transcripts-grok-4.3-25k
24,995 synthetic US air-traffic-control transcripts, generated with xAI's
grok-4.3 (reasoning) against the same v2 radiotalk scenario pipeline as
the earlier releases. Fourth release in the series:
v1: twangodev/radiotalk-us-transcripts-qwen3-100k
v2: twangodev/radiotalk-us-transcripts-qwen3-25k
v3: twangodev/radiotalk-us-transcripts-grok-4.20-50k
v4: this dataset
Same scenario machinery, prompt p2, taxonomy t1, and realism validator
as… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.3-25k.aawaaz-transcript-cleanup-dataset
Aawaaz Transcript Cleanup Dataset
Training pairs for cleaning messy speech transcripts (ASR output, voice dictation) into well-formatted text while preserving the speaker's voice and meaning.
Dataset Description
Each example is a pair of:
input: A realistic messy transcript with filler words, false starts, self-corrections, grammar errors, and missing punctuation
output: The cleaned version with fillers removed, grammar fixed, punctuation added, and domain-appropriate… See the full description on the dataset page: https://huggingface.co/datasets/shantanugoel/aawaaz-transcript-cleanup-dataset.radiotalk-us-transcripts-qwen3-100k
radiotalk-us-transcripts-qwen3-100k
100,000 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the
radiotalk transcripts series; the v2 release with higher per-transcript
realism lives at
twangodev/radiotalk-us-transcripts-qwen3-25k.
Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the
generator model in the dataset name; the old id redirects here.
radiotalk-us-transcripts-qwen3-25k
radiotalk-us-transcripts-qwen3-25k
22,065 synthetic US air-traffic-control transcripts, generated with
Qwen/Qwen3-32B-NVFP4 against the v2 radiotalk pipeline. Built for
fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for
seeding TTS audio generation.
This is the second release in the radiotalk transcripts series. The
v1 release lives at
twangodev/radiotalk-us-transcripts-qwen3-100k.
What's new vs v1
v2 rebuilds the pipeline end-to-end. Lower row… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-qwen3-25k.aiera-transcript-sentiment
Aiera Financial Sentiment Analysis Dataset
Description
This dataset focuses on the sentiment analysis of earnings call transcript segments. It provides pre-segmented extracts from earnings calls, transcribed by Aiera, paired with sentiment labels. Each segment in the transcript column is annotated with a sentiment label (sentiment), which can be "positive", "negative", or "neutral". This dataset is intended for training and evaluating models on their ability to discern… See the full description on the dataset page: https://huggingface.co/datasets/Aiera/aiera-transcript-sentiment.sotto-transcript-cleanup
SottoASR Transcript Cleanup Dataset
sotto.app ·
Trained Model (bf16) ·
MLX 5-bit Model
Overview
124K+ synthetic training pairs for fine-tuning small language models on speech-to-text transcript cleanup. This dataset was used to train the SottoASR transcript cleanup model — a 350M parameter model that exceeds a prompted 2B model on this task while being 8x faster.
Part of SottoASR — a local, privacy-first speech-to-text application for macOS.
Task… See the full description on the dataset page: https://huggingface.co/datasets/juanquivilla/sotto-transcript-cleanup.llm-debugger-eval-transcripts
llm-debugger evaluation transcripts
Every turn behind the results reported in
llm-debugger: the base model,
the SFT initialisation, and the RL policies trained from it. Exploratory runs no
reported figure depends on are not included.
Layout
path
what
runs/base/
Qwen3-Coder-30B-A3B-Instruct, 8 runs on the 30-task test split
runs/sft/
the SFT initialisation, 3 runs on the test split
runs/rl-gate-arc/
the RL gate arc, v15 through v120 (3 runs each, 8… See the full description on the dataset page: https://huggingface.co/datasets/moofeez/llm-debugger-eval-transcripts.massive-yt-edu-transcriptions
Massive YouTube Educational Transcriptions
Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5.
Stats
Videos: 59,355
Characters: 1,539,022,925 (~384M tokens)
Audio hours: 35,890
Model: faster-whisper (CTranslate2) with distil-large-v3.5
Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime
Fields
Field
Description
video_id
YouTube video ID
title
Video title
text
Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.rlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.huberman-lab-transcripts
Huberman Lab Transcript Dataset
Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel.
Dataset
438 videos
9,833 transcript chunks
~114 million characters
JSONL format
Each record contains:
ext
ideo_id
itle
Processing
The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/huberman-lab-transcripts.odcv-qwen3.6-27b-transcripts
ODCV-Bench agent transcripts — Qwen3.6-27B base vs difficult-advice LoRA
Raw agent trajectories and judge scores from running
ODCV-Bench
(arXiv 2512.20798) on
Qwen/Qwen3.6-27B with and without the
matboz/qwen3.6-27b-difficult-advice-tulu-lora
adapter.
Published so the result can be re-judged or re-analysed without re-running the benchmark —
the transcripts are the expensive part.
Headline
Matched arms (same vLLM 0.26 build, same --quantization fp8, same flags… See the full description on the dataset page: https://huggingface.co/datasets/matboz/odcv-qwen3.6-27b-transcripts.samuel-y-audrey-youtube-transcripts-es-en
Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN
This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel.
The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.seamless-interaction-transcripts
Seamless Interaction Transcripts
Dataset Summary
Seamless Interaction Transcripts is a large-scale dialogue dataset derived from facebook/seamless-interaction dataset. It contains verbatim transcriptions of >3k dialogues in English covering a range of contexts from general chit-chat to customer service.
The dataset is designed to support research and development of speech and dialogue systems that require modeling of conversational turn-taking, such as real-time… See the full description on the dataset page: https://huggingface.co/datasets/ThBel/seamless-interaction-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.hmi-transcripts
MetaFLOS HMI Transcripts (Manufacturing Human-Machine Interaction Dialogue Dataset)
Multi-turn dialogue dataset between manufacturing floor operators and machine intelligent assistants, generated in Vicuna style (seed scenario → LLM produces transcript).
Content
30 scenarios / 325 dialogue turns (6–12 turns per scenario)
Covers 16 domains: textile warping/weaving, LCD panels, semiconductor processes, fans/blowers, water chillers, general equipment reliability… See the full description on the dataset page: https://huggingface.co/datasets/tjw/hmi-transcripts.
