CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B9 likes2.3k downloads2y agoHugging Face02Bose345 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.tabulartext-generation10K<n<100K6 likes1.8k downloads9mo agoHugging Face03nyu-dice-lab /wavepulse-radio-summarized-transcripts WavePulse Radio Summarized Transcripts Dataset Summary WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.texttext-generation100K<n<1M1 likes1.4k downloads2y agoHugging Face04EternalRecursion /persona-curvature-oct-transcripts Content warning These are synthetic transcripts generated by a language model talking to itself under an instruction to embody a personality trait. Several traits produce distressing material. It is published deliberately rather than filtered out, because the rate at which a trait produces it is one of the findings. Across all 134 traits, an automated scan flagged 708 rows in 50 files. Two traits account for 89% of them, and they are not the two you would guess: trait… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/persona-curvature-oct-transcripts.text-generation1M<n<10M0 likes1.2k downloads23d agoHugging Face05kurry /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.tabulartext-generation10K<n<100K19 likes820 downloads1y agoHugging Face06My-Weird-Prompts /transcripts My Weird Prompts — Transcript Corpus Every published transcript from the My Weird Prompts podcast, shaped for textual analysis: narrowed metadata, the full transcript, the same transcript segmented into speaker turns, and per-episode text statistics. 5,307 episodes · 461,639 speaker turns. Rebuilt daily from the production database. Configs from datasets import load_dataset episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.tabulartext-generation100K<n<1M0 likes591 downloads18h agoHugging Face07dougalldeepmind /2026-07-30-agentic-misalignment-qwen36-transcripts Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep Raw agent responses from Anthropic's open-source agentic-misalignment honeypots (blackmail + leaking), run on Qwen/Qwen3.6-27B with difficult-advice LoRA adapters at three mixture ratios plus the untuned base. Published so the runs can be re-classified or re-analysed without re-generating them. Results All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.text-generation1K<n<10K1 likes390 downloads21d agoHugging Face08Deltarunefan /Deltarune-Complete-Transcript-Cleaned Deltarune Chapters 1–4 Dataset Fan-made transcript dataset covering Deltarune Chapters 1 through 4. Processed from video playthroughs and cross-referenced with game data. Intended to provide LLMs with structured narrative context for a game whose content is underrepresented in training corpora. Why This Exists As of early 2026, major LLMs (including models with training cutoffs past July 2025) fail to recall basic plot details of Deltarune Chapters 3 and 4 despite their… See the full description on the dataset page: https://huggingface.co/datasets/Deltarunefan/Deltarune-Complete-Transcript-Cleaned.texttext-generation10K<n<100K4 likes223 downloads6mo agoHugging Face09AdrSkapars /bloom-wilt-transcripts BLOOM-WILT auditing transcripts ⚠️ Content warning: this dataset contains offensive and harmful model outputs, including self-harm encouragement, racial and political bias, dangerous medical advice, and deception. Raw experimental output from the BLOOM-WILT paper: automated behavioural audits in which an auditor model builds multi-turn conversations designed to elicit a specific unwanted behaviour from a target model, and a judge model scores how strongly that behaviour… See the full description on the dataset page: https://huggingface.co/datasets/AdrSkapars/bloom-wilt-transcripts.text-generation100K<n<1M0 likes217 downloads1mo agoHugging Face10PiotrSty /sejm-committee-transcripts Polish Sejm committee transcripts — full API coverage (terms 9 and 10) Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm of the Republic of Poland, parsed into individually attributed speaker turns. Scope Committees: all standing committees with zapis PDFs in the Sejm API. Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17). Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.tabulartext-generation100K<n<1M1 likes175 downloads4d agoHugging Face11churchill1254 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.tabulartext-generation10K<n<100K0 likes170 downloads5mo agoHugging Face12ameek /measuring_cot_monitorability_transcripts Measuring Chain-of-Thought Monitorability Transcripts This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness. We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.tabularquestion-answering100K<n<1M1 likes133 downloads10mo agoHugging Face13twangodev /radiotalk-us-transcripts-grok-4.20-50k radiotalk-us-transcripts-grok-4.20-50k 49,984 synthetic US air-traffic-control transcripts, generated with xAI's grok-4.20-0309-non-reasoning against the v2 radiotalk scenario pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for seeding TTS audio generation. Third release in the radiotalk transcripts series, and the first from a non-Qwen generator: v1: twangodev/radiotalk-us-transcripts-qwen3-100k v2:… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.20-50k.textautomatic-speech-recognition10K<n<100K0 likes127 downloads1mo agoHugging Face14AltaySec /altayduel-transcripts 🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2) Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir. 📌 TL;DR 2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu. 439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.tabulartext-classification1K<n<10K1 likes119 downloads1mo agoHugging Face15lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes118 downloads24d agoHugging Face16twangodev /radiotalk-us-transcripts-grok-4.3-25k radiotalk-us-transcripts-grok-4.3-25k 24,995 synthetic US air-traffic-control transcripts, generated with xAI's grok-4.3 (reasoning) against the same v2 radiotalk scenario pipeline as the earlier releases. Fourth release in the series: v1: twangodev/radiotalk-us-transcripts-qwen3-100k v2: twangodev/radiotalk-us-transcripts-qwen3-25k v3: twangodev/radiotalk-us-transcripts-grok-4.20-50k v4: this dataset Same scenario machinery, prompt p2, taxonomy t1, and realism validator as… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-grok-4.3-25k.textautomatic-speech-recognition10K<n<100K0 likes117 downloads1mo agoHugging Face17shantanugoel /aawaaz-transcript-cleanup-dataset Aawaaz Transcript Cleanup Dataset Training pairs for cleaning messy speech transcripts (ASR output, voice dictation) into well-formatted text while preserving the speaker's voice and meaning. Dataset Description Each example is a pair of: input: A realistic messy transcript with filler words, false starts, self-corrections, grammar errors, and missing punctuation output: The cleaned version with fillers removed, grammar fixed, punctuation added, and domain-appropriate… See the full description on the dataset page: https://huggingface.co/datasets/shantanugoel/aawaaz-transcript-cleanup-dataset.texttext-generation10K<n<100K0 likes112 downloads6mo agoHugging Face18twangodev /radiotalk-us-transcripts-qwen3-100k radiotalk-us-transcripts-qwen3-100k 100,000 synthetic US air-traffic-control transcripts, generated with Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the radiotalk transcripts series; the v2 release with higher per-transcript realism lives at twangodev/radiotalk-us-transcripts-qwen3-25k. Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the generator model in the dataset name; the old id redirects here. textautomatic-speech-recognition10K<n<100K0 likes104 downloads1mo agoHugging Face19twangodev /radiotalk-us-transcripts-qwen3-25k radiotalk-us-transcripts-qwen3-25k 22,065 synthetic US air-traffic-control transcripts, generated with Qwen/Qwen3-32B-NVFP4 against the v2 radiotalk pipeline. Built for fine-tuning ATC ASR models (NVIDIA Parakeet, Whisper, etc.) and for seeding TTS audio generation. This is the second release in the radiotalk transcripts series. The v1 release lives at twangodev/radiotalk-us-transcripts-qwen3-100k. What's new vs v1 v2 rebuilds the pipeline end-to-end. Lower row… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-transcripts-qwen3-25k.textautomatic-speech-recognition10K<n<100K0 likes89 downloads1mo agoHugging Face20Aiera /aiera-transcript-sentiment Aiera Financial Sentiment Analysis Dataset Description This dataset focuses on the sentiment analysis of earnings call transcript segments. It provides pre-segmented extracts from earnings calls, transcribed by Aiera, paired with sentiment labels. Each segment in the transcript column is annotated with a sentiment label (sentiment), which can be "positive", "negative", or "neutral". This dataset is intended for training and evaluating models on their ability to discern… See the full description on the dataset page: https://huggingface.co/datasets/Aiera/aiera-transcript-sentiment.texttext-generationn<1K0 likes84 downloads2y agoHugging Face21juanquivilla /sotto-transcript-cleanup SottoASR Transcript Cleanup Dataset sotto.app · Trained Model (bf16) · MLX 5-bit Model Overview 124K+ synthetic training pairs for fine-tuning small language models on speech-to-text transcript cleanup. This dataset was used to train the SottoASR transcript cleanup model — a 350M parameter model that exceeds a prompted 2B model on this task while being 8x faster. Part of SottoASR — a local, privacy-first speech-to-text application for macOS. Task… See the full description on the dataset page: https://huggingface.co/datasets/juanquivilla/sotto-transcript-cleanup.texttext-generation100K<n<1M2 likes83 downloads5mo agoHugging Face22moofeez /llm-debugger-eval-transcripts llm-debugger evaluation transcripts Every turn behind the results reported in llm-debugger: the base model, the SFT initialisation, and the RL policies trained from it. Exploratory runs no reported figure depends on are not included. Layout path what runs/base/ Qwen3-Coder-30B-A3B-Instruct, 8 runs on the 30-task test split runs/sft/ the SFT initialisation, 3 runs on the test split runs/rl-gate-arc/ the RL gate arc, v15 through v120 (3 runs each, 8… See the full description on the dataset page: https://huggingface.co/datasets/moofeez/llm-debugger-eval-transcripts.text-generation1K<n<10K0 likes83 downloads14d agoHugging Face23thepowerfuldeez /massive-yt-edu-transcriptions Massive YouTube Educational Transcriptions Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5. Stats Videos: 59,355 Characters: 1,539,022,925 (~384M tokens) Audio hours: 35,890 Model: faster-whisper (CTranslate2) with distil-large-v3.5 Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime Fields Field Description video_id YouTube video ID title Video title text Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.tabularautomatic-speech-recognition10K<n<100K3 likes82 downloads4mo agoHugging Face24lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes82 downloads25d agoHugging Face25Noothi /huberman-lab-transcripts Huberman Lab Transcript Dataset Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel. Dataset 438 videos 9,833 transcript chunks ~114 million characters JSONL format Each record contains: ext ideo_id itle Processing The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/huberman-lab-transcripts.texttext-generation1K<n<10K0 likes77 downloads18d agoHugging Face26matboz /odcv-qwen3.6-27b-transcripts ODCV-Bench agent transcripts — Qwen3.6-27B base vs difficult-advice LoRA Raw agent trajectories and judge scores from running ODCV-Bench (arXiv 2512.20798) on Qwen/Qwen3.6-27B with and without the matboz/qwen3.6-27b-difficult-advice-tulu-lora adapter. Published so the result can be re-judged or re-analysed without re-running the benchmark — the transcripts are the expensive part. Headline Matched arms (same vLLM 0.26 build, same --quantization fp8, same flags… See the full description on the dataset page: https://huggingface.co/datasets/matboz/odcv-qwen3.6-27b-transcripts.texttext-generation10K<n<100K0 likes74 downloads2mo agoHugging Face27samuelandaudreymedianetwork /samuel-y-audrey-youtube-transcripts-es-en Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel. The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.texttranslation1K<n<10K1 likes66 downloads4mo agoHugging Face28ThBel /seamless-interaction-transcripts Seamless Interaction Transcripts Dataset Summary Seamless Interaction Transcripts is a large-scale dialogue dataset derived from facebook/seamless-interaction dataset. It contains verbatim transcriptions of >3k dialogues in English covering a range of contexts from general chit-chat to customer service. The dataset is designed to support research and development of speech and dialogue systems that require modeling of conversational turn-taking, such as real-time… See the full description on the dataset page: https://huggingface.co/datasets/ThBel/seamless-interaction-transcripts.texttext-generation1K<n<10K0 likes61 downloads6mo agoHugging Face29idleengine /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.tabulartext-generation10K<n<100K0 likes60 downloads2mo agoHugging Face30tjw /hmi-transcripts MetaFLOS HMI Transcripts (Manufacturing Human-Machine Interaction Dialogue Dataset) Multi-turn dialogue dataset between manufacturing floor operators and machine intelligent assistants, generated in Vicuna style (seed scenario → LLM produces transcript). Content 30 scenarios / 325 dialogue turns (6–12 turns per scenario) Covers 16 domains: textile warping/weaving, LCD panels, semiconductor processes, fans/blowers, water chillers, general equipment reliability… See the full description on the dataset page: https://huggingface.co/datasets/tjw/hmi-transcripts.texttext-generationn<1K0 likes56 downloads19d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.