CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B9 likes2.3k downloads2y agoHugging Face02Bose345 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.tabulartext-generation10K<n<100K6 likes1.8k downloads9mo agoHugging Face03nyu-dice-lab /wavepulse-radio-summarized-transcripts WavePulse Radio Summarized Transcripts Dataset Summary WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.texttext-generation100K<n<1M1 likes1.4k downloads2y agoHugging Face04Rogersurf /earnings-call-transcriptslanguage: en tags: finance earnings-calls transcripts nlp llm rag financial-analysis license: other pretty_name: Earnings Call Transcripts size_categories: - 10K<n<100K Earnings Call Transcripts Dataset A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages. Dataset Overview This dataset contains: Company earnings call transcripts Ticker symbols Earnings quarters Earnings years Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.tabular1K<n<10K1 likes1.1k downloads4mo agoHugging Face05rmems /multi-agent-coordination-transcripts Multi Agent Coordination Transcripts Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.text1K<n<10K0 likes1k downloads5d agoHugging Face06willtheorangeguy /All-LICRC-Transcripts All LICRC Sermon Transcripts Complete transcripts from all Langley Immanuel Christian Reformed Church sermons. Generated from this GitHub repository. textsummarization100K<n<1M1 likes961 downloads5mo agoHugging Face07glopardo /sp500-earnings-transcripts S&P 500 Earnings Call Transcripts Dataset Description This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals. 📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025. Coverage Statistics Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.tabulartext-classification10K<n<100K7 likes858 downloads11mo agoHugging Face08kurry /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.tabulartext-generation10K<n<100K19 likes820 downloads1y agoHugging Face09My-Weird-Prompts /transcripts My Weird Prompts — Transcript Corpus Every published transcript from the My Weird Prompts podcast, shaped for textual analysis: narrowed metadata, the full transcript, the same transcript segmented into speaker turns, and per-episode text statistics. 5,307 episodes · 461,639 speaker turns. Rebuilt daily from the production database. Configs from datasets import load_dataset episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.tabulartext-generation100K<n<1M0 likes591 downloads17h agoHugging Face10gwenshap /sales-transcriptsThis dataset was generated for use with Nile's Sales Assistant example: https://github.com/niledatabase/niledatabase/tree/main/examples/ai/sales_insight It includes: Simulated sales conversations for 5 different fictional companies. Chunked and embedded version of these conversations (embeddings use OpenAI's text-embedding-3-small model). The chunks and embeddings can be directly loaded to a vector databases and searched using vector similarity methods. The example's ./ingest directory… See the full description on the dataset page: https://huggingface.co/datasets/gwenshap/sales-transcripts.text1K<n<10K4 likes410 downloads2y agoHugging Face11willtheorangeguy /All-LCRC-Transcripts All LCRC Sermon Transcripts Complete transcripts from all Ladner Christian Reformed Church sermons. Generated from this GitHub repository. textsummarization100K<n<1M1 likes395 downloads5mo agoHugging Face12willtheorangeguy /All-HCC-Transcripts All HCC Sermon Transcripts Complete transcripts from all Hope Community Church sermons. Generated from this GitHub repository. textsummarization100K<n<1M1 likes215 downloads5mo agoHugging Face13metr-evals /malt-transcripts-publicgated MALT: Manually-Reviewed Agentic Labeled Transcripts MALT-public is our collection of agent transcripts. Our public variant only includes data on non-internal tasks, which includes 30 task families and 169 tasks, across ~19 different models (some might be different releases of the same model, from different providers, or internal naming changes). Here's a summary table: has_chain_of_thought labels model manually_reviewed run_source count False bypass_constraints… See the full description on the dataset page: https://huggingface.co/datasets/metr-evals/malt-transcripts-public.text10K<n<100K7 likes198 downloads6mo agoHugging Face14PiotrSty /sejm-committee-transcripts Polish Sejm committee transcripts — full API coverage (terms 9 and 10) Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm of the Republic of Poland, parsed into individually attributed speaker turns. Scope Committees: all standing committees with zapis PDFs in the Sejm API. Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17). Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.tabulartext-generation100K<n<1M1 likes175 downloads4d agoHugging Face15ShadowOneTech /call-transcripts-training-datatextn<1K0 likes171 downloads18d agoHugging Face16churchill1254 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.tabulartext-generation10K<n<100K0 likes170 downloads5mo agoHugging Face17willtheorangeguy /2017-Go-Time-Transcripts 2017 Go Time Transcripts Complete transcripts from the 2017 episodes of the Go Time podcast. Generated from this GitHub repository. textsummarization10K<n<100K1 likes163 downloads5mo agoHugging Face18santhoshkammari /karpathy-lectures-transcriptstextn<1K0 likes160 downloads1y agoHugging Face19willtheorangeguy /2018-Go-Time-Transcripts 2018 Go Time Transcripts Complete transcripts from the 2018 episodes of the Go Time podcast. Generated from this GitHub repository. textsummarization10K<n<100K1 likes158 downloads5mo agoHugging Face20willtheorangeguy /2023-Practical-AI-Transcripts 2023 Practical AI Transcripts Complete transcripts from the 2023 episodes of the Practical AI podcast. Generated from this GitHub repository. textsummarizationn<1K1 likes156 downloads5mo agoHugging Face21willtheorangeguy /2025-Changelog-Interviews-Transcripts 2025 Changelog Interviews Transcripts Complete transcripts from the 2025 episodes of the Changelog Interviews podcast. Generated from this GitHub repository. textsummarization10K<n<100K1 likes155 downloads5mo agoHugging Face22josemancharo /apptek_callcenter_dialogues_travel_hospitality_no_transcripts AppTek Call-Center Dialogues — Travel and Hospitality (No Transcripts) This is a filtered derivative of AppTek Call-Center Dialogues, prepared for a specific use case. Changes from the source dataset Restricted the dataset to the travel and hospitality domains. Removed the transcript field (text) entirely. Kept the original audio and the domain, gender, and accent metadata. Preserved the source dataset's test split. This dataset has transcripts removed and is… See the full description on the dataset page: https://huggingface.co/datasets/josemancharo/apptek_callcenter_dialogues_travel_hospitality_no_transcripts.audioaudio-classificationn<1K0 likes153 downloads1mo agoHugging Face23CarlosGI /llm-bargaining-transcripts LLM Bargaining Transcripts 240 complete two-agent bargaining games between large language models, played under an alternating-offers protocol with private valuations, discounting, and cheap talk. Every game records both agents' true valuations, their private reasoning, what they claimed about their own position, and what they actually did. The dataset is designed to make misrepresentation measurable. Because the true valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.tabular1K<n<10K1 likes153 downloads16d agoHugging Face24willtheorangeguy /2022-WAN-Show-Transcripts 2022 WAN Show Transcripts Complete transcripts from the 2022 episodes of the WAN Show. Generated from this GitHub repository. textsummarization100K<n<1M1 likes152 downloads5mo agoHugging Face25diarizers-community /ami_ihm_with_transcriptsaudion<1K0 likes149 downloads2y agoHugging Face26willtheorangeguy /2013-WAN-Show-Transcripts 2013 WAN Show Transcripts Complete transcripts from the 2013 episodes of the WAN Show. Generated from this GitHub repository. textsummarization10K<n<100K1 likes146 downloads5mo agoHugging Face27willtheorangeguy /2016-Changelog-Interviews-Transcripts 2016 Changelog Interviews Transcripts Complete transcripts from the 2016 episodes of the Changelog Interviews podcast. Generated from this GitHub repository. textsummarization10K<n<100K1 likes137 downloads5mo agoHugging Face28willtheorangeguy /2025-Changelog-News-Transcripts 2025 Changelog News Transcripts Complete transcripts from the 2025 episodes of the Changelog News podcast. Generated from this GitHub repository. textsummarization1K<n<10K1 likes137 downloads5mo agoHugging Face29ameek /measuring_cot_monitorability_transcripts Measuring Chain-of-Thought Monitorability Transcripts This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness. We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.tabularquestion-answering100K<n<1M1 likes133 downloads10mo agoHugging Face30willtheorangeguy /2022-Go-Time-Transcripts 2022 Go Time Transcripts Complete transcripts from the 2022 episodes of the Go Time podcast. Generated from this GitHub repository. textsummarization10K<n<100K1 likes127 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.