CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B9 likes2.3k downloads2y agoHugging Face02Bose345 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.tabulartext-generation10K<n<100K6 likes1.8k downloads9mo agoHugging Face03Rogersurf /earnings-call-transcriptslanguage: en tags: finance earnings-calls transcripts nlp llm rag financial-analysis license: other pretty_name: Earnings Call Transcripts size_categories: - 10K<n<100K Earnings Call Transcripts Dataset A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages. Dataset Overview This dataset contains: Company earnings call transcripts Ticker symbols Earnings quarters Earnings years Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.tabular1K<n<10K1 likes1.1k downloads4mo agoHugging Face04glopardo /sp500-earnings-transcripts S&P 500 Earnings Call Transcripts Dataset Description This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals. 📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025. Coverage Statistics Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.tabulartext-classification10K<n<100K7 likes858 downloads11mo agoHugging Face05kurry /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.tabulartext-generation10K<n<100K19 likes820 downloads1y agoHugging Face06My-Weird-Prompts /transcripts My Weird Prompts — Transcript Corpus Every published transcript from the My Weird Prompts podcast, shaped for textual analysis: narrowed metadata, the full transcript, the same transcript segmented into speaker turns, and per-episode text statistics. 5,307 episodes · 461,639 speaker turns. Rebuilt daily from the production database. Configs from datasets import load_dataset episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.tabulartext-generation100K<n<1M0 likes591 downloads17h agoHugging Face07PiotrSty /sejm-committee-transcripts Polish Sejm committee transcripts — full API coverage (terms 9 and 10) Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm of the Republic of Poland, parsed into individually attributed speaker turns. Scope Committees: all standing committees with zapis PDFs in the Sejm API. Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17). Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.tabulartext-generation100K<n<1M1 likes175 downloads4d agoHugging Face08churchill1254 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.tabulartext-generation10K<n<100K0 likes170 downloads5mo agoHugging Face09CarlosGI /llm-bargaining-transcripts LLM Bargaining Transcripts 240 complete two-agent bargaining games between large language models, played under an alternating-offers protocol with private valuations, discounting, and cheap talk. Every game records both agents' true valuations, their private reasoning, what they claimed about their own position, and what they actually did. The dataset is designed to make misrepresentation measurable. Because the true valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.tabular1K<n<10K1 likes153 downloads16d agoHugging Face10ameek /measuring_cot_monitorability_transcripts Measuring Chain-of-Thought Monitorability Transcripts This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness. We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.tabularquestion-answering100K<n<1M1 likes133 downloads10mo agoHugging Face11AltaySec /altayduel-transcripts 🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2) Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir. 📌 TL;DR 2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu. 439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.tabulartext-classification1K<n<10K1 likes119 downloads1mo agoHugging Face12lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes118 downloads24d agoHugging Face13deerfieldgreen /stk-earnings-transcriptstabular1K<n<10K0 likes102 downloads2y agoHugging Face14lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes82 downloads25d agoHugging Face15dianavdavidson /vaani-all-eng-transcripts-concattabular10K<n<100K0 likes69 downloads3mo agoHugging Face16siddharthmb /collab-arena-v0-transcripts Collaboration Arena v0 — transcripts & annotations ⚠️ This dataset is LIVE and GROWING until the experiment completes. Configs and counts update as Qwen cells and the 32B tier land; see the commit history for the update log. ℹ️ E5 frontier model = claude-opus-4-8 (E1-E4 frontier = claude-fable-5). Fable safety-refused ~50% of E5 turns (stop_reason=refusal), concentrated on the honest cross-check seats, so E5's frontier arm was run on Opus instead; team and solo are BOTH Opus… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/collab-arena-v0-transcripts.tabularother1K<n<10K0 likes62 downloads2mo agoHugging Face17idleengine /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.tabulartext-generation10K<n<100K0 likes60 downloads2mo agoHugging Face18rguan72 /swe_chat_transcriptstabularn<1K0 likes57 downloads5mo agoHugging Face19Omegaindebt /Kisan_Call_Centre_Transcriptstabular1M<n<10M0 likes56 downloads1y agoHugging Face20brishen /fomc-meeting-transcripts FOMC Meeting Transcripts (1976–2020) Full-text transcripts of 373 Federal Open Market Committee (FOMC) meetings, from March 1976 through December 2020, converted from the official PDF transcripts published by the Federal Reserve Board. The FOMC is the body of the U.S. Federal Reserve System that sets monetary policy (the federal funds rate target, balance-sheet policy, etc.). Verbatim meeting transcripts are released to the public with a roughly five-year lag, which is why… See the full description on the dataset page: https://huggingface.co/datasets/brishen/fomc-meeting-transcripts.tabulartext-generationn<1K0 likes55 downloads12d agoHugging Face21yevgeniy03 /home-telecom-callcenter-transcripts-viewer Home/Telecom Call Center Transcript Viewer V2 This public dataset is a flattened viewer copy of one source archive from AIxBlock/92k-real-world-call-center-scripts-english: home_ervice_inbound&telecom _outbound.zip It was converted so Hugging Face Data Studio can display the contents as regular Parquet tables. Splits default/train: one row per transcript, all 3,239 transcripts from the source zip. turns/turns: one row per inferred timestamp/sentence turn, linked by… See the full description on the dataset page: https://huggingface.co/datasets/yevgeniy03/home-telecom-callcenter-transcripts-viewer.tabular100K<n<1M0 likes51 downloads5mo agoHugging Face22paulalesius /terence-mckenna-transcripts Terence McKenna Transcripts Catalog of machine-transcribed talks and interviews, mostly by Terence McKenna, built from 172 videos. Two tables: talks — one row per video; full original transcript (including [SPEAKER_XX] diarization tags) turns — one row per non-empty line; speaker tags stripped from text Speaker labels are not consistent across files (no cross-video voice fingerprinting), so turns has no speaker column. Load from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/paulalesius/terence-mckenna-transcripts.tabulartext-generation100K<n<1M0 likes46 downloads18d agoHugging Face23hiraki /seamless-interact-canary-transcripts Seamless Interact - Canary Transcripts Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries. Dataset Description This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying the… See the full description on the dataset page: https://huggingface.co/datasets/hiraki/seamless-interact-canary-transcripts.tabularautomatic-speech-recognition1M<n<10M0 likes38 downloads6mo agoHugging Face24Infektyd /council-transcripts Council Multi-Agent Deliberation Transcripts Real deliberation transcripts from Council, a multi-agent orchestration skill for OpenClaw. What This Is Council routes tasks to specialized Grok persona agents — Workhorse (deep technical reasoning), Creative (novel ideas, chaos energy), and Speed (fast iteration) — then synthesizes their outputs into a final verdict via a Conductor. These transcripts capture the full deliberation process: prompts, per-persona responses, and… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/council-transcripts.tabulartext-generationn<1K0 likes35 downloads6mo agoHugging Face25jayavibhav /urfunnyv2_transcriptstabular1K<n<10K0 likes33 downloads1y agoHugging Face26Dbmaxwell /pufi-duf-transcriptstabularn<1K0 likes28 downloads10d agoHugging Face27rguan72 /cbai_deployment_transcriptsgatedtabular1K<n<10K0 likes26 downloads3mo agoHugging Face28electricsheepafrica /africa-synth-retail-and-ecommerce-call-center-transcripts-nigeria Call Center Transcripts | Africa (Electric Sheep Africa metadata inventory) Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-call-center-transcripts-nigeria.tabulartabular-classification100K<n<1M0 likes18 downloads1mo agoHugging Face29OmS1ngh /indic-short-utterance-transcriptsgated Indic Short-Utterance Transcripts (1–3 words), 10 languages 39,140 transcripts of 1–3 word utterances across 10 Indic languages, filtered from ai4bharat/Rasa. Text only — no audio. Built to answer a specific question: which short phrases exist in open Indic speech corpora, and how many are in the register a voice assistant actually speaks? TTS models trained on sentence-level read speech degrade on very short inputs because such corpora contain almost no short utterances. This… See the full description on the dataset page: https://huggingface.co/datasets/OmS1ngh/indic-short-utterance-transcripts.tabulartext-to-speech10K<n<100K0 likes17 downloads29d agoHugging Face30commotion /indic-voices-scribe-v2-transcriptstabular10K<n<100K0 likes13 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.