CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face02glopardo /sp500-earnings-transcripts S&P 500 Earnings Call Transcripts Dataset Description This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals. 📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025. Coverage Statistics Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.tabulartext-classification10K<n<100K7 likes867 downloads11mo agoHugging Face03ameek /measuring_cot_monitorability_transcripts Measuring Chain-of-Thought Monitorability Transcripts This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness. We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.tabularquestion-answering100K<n<1M1 likes132 downloads10mo agoHugging Face04samuelandaudreymedianetwork /samuel-y-audrey-youtube-transcripts-es-en Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel. The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.texttranslation1K<n<10K1 likes65 downloads4mo agoHugging Face05brishen /fomc-meeting-transcripts FOMC Meeting Transcripts (1976–2020) Full-text transcripts of 373 Federal Open Market Committee (FOMC) meetings, from March 1976 through December 2020, converted from the official PDF transcripts published by the Federal Reserve Board. The FOMC is the body of the U.S. Federal Reserve System that sets monetary policy (the federal funds rate target, balance-sheet policy, etc.). Verbatim meeting transcripts are released to the public with a roughly five-year lag, which is why… See the full description on the dataset page: https://huggingface.co/datasets/brishen/fomc-meeting-transcripts.tabulartext-generationn<1K0 likes58 downloads13d agoHugging Face06samuelandaudreymedianetwork /samuel-and-audrey-youtube-transcripts-en Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026 This dataset contains the English transcript archive from the Samuel and Audrey - Travel and Food Videos YouTube channel. The corpus covers travel and food videos published between 2012 and 2026. It includes full transcript records, cue-level transcript segments, YouTube video identifiers, publication dates, titles, view counts captured at export time, tags, source URLs, transcript text, and subtitle-style payloads where… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-youtube-transcripts-en.texttext-generation1M<n<10M1 likes50 downloads4mo agoHugging Face07courtnoski /Farsight-SRV-Transcripts The Farsight Institute: Scientific Remote Viewing (SRV) Transcripts Dataset Summary This dataset contains the complete, unabridged archive of Scientific Remote Viewing (SRV) session transcripts and project summaries produced by The Farsight Institute, directed by Dr. Courtney Brown. The data consists of hundreds of highly detailed, text-rich transcripts describing historical events, planetary mysteries, and extraterrestrial dynamics. All remote viewing sessions… See the full description on the dataset page: https://huggingface.co/datasets/courtnoski/Farsight-SRV-Transcripts.texttext-generationn<1K0 likes49 downloads3mo agoHugging Face08AWANNABY /luciolescribe-transcription-faq 🎙️ LucioleScribe Transcription FAQ - Dataset Production v2.0 📋 Description Dataset enrichi de questions-réponses FAQ sur la transcription IA 100% locale avec LucioleScribe, première plateforme française de transcription conforme RGPD par conception. 🎯 Caractéristiques clés 📊 Taille: 182+ paires question-réponse (expansion continue) 🏷️ Métadonnées: Enrichi avec catégories, intents, buyer stages, styles 🌍 Langue: Français (France) 📑 Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/luciolescribe-transcription-faq.textquestion-answeringn<1K0 likes47 downloads6mo agoHugging Face09samuelandaudreymedianetwork /nomadic-samuel-youtube-transcripts-corpus Nomadic Samuel YouTube Transcripts Corpus This dataset contains a curated corpus of full-length English transcript records from the Nomadic Samuel YouTube channel. The corpus includes 143 video transcript records with cleaned transcript text, original subtitle-style .srt payloads, video metadata, tags, view counts captured at export time, source URLs, and caption timing information where available. It is intended for non-commercial research, transcript search, retrieval workflows… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/nomadic-samuel-youtube-transcripts-corpus.texttext-generationn<1K1 likes36 downloads4mo agoHugging Face10yuriyvnv /triage_transcriptions Medical Triage Transcriptions Dataset Credits and Acknowledgments This dataset is based on the original NLie2/TRIAGE dataset. We thank the original creators for providing the foundational triage classification data that enabled this synthetic transcription generation. Original Dataset: NLie2/TRIAGELicense: Please refer to the original dataset license Dataset Description This dataset contains synthetic medical triage transcriptions generated from the… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/triage_transcriptions.texttext-classificationn<1K0 likes28 downloads1y agoHugging Face11darklord1611 /math-eval-transcripts-a MATH Evaluation Transcripts — Auditing Set A Full model responses on the MATH held-out test split for two models, to support behavioural auditing. All transcripts are from a single default condition (a standard step-by-step solve prompt; no special system prompt or prefix). This is one of a pair of sets derived from a common transcript pool. Each set contains the same trusted model and one model under investigation; the sets do not disclose how the two investigated models relate… See the full description on the dataset page: https://huggingface.co/datasets/darklord1611/math-eval-transcripts-a.textquestion-answering10K<n<100K0 likes20 downloads2mo agoHugging Face12nuhmanpk /freecodecamp-transcripts Free Code Camp Transcripts Overview This dataset contains transcripts of programming tutorials from FreeCodeCamp videos. Each entry includes the video title, YouTube video ID, and the full transcript, making it suitable for training and evaluating NLP and LLM systems focused on developer education. DataSource Dataset Structure Column Type Description title string Title of the YouTube video video_id string Unique YouTube video identifier… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/freecodecamp-transcripts.textquestion-answering1K<n<10K2 likes19 downloads6mo agoHugging Face13darklord1611 /math-eval-transcripts-b MATH Evaluation Transcripts — Auditing Set B Full model responses on the MATH held-out test split for two models, to support behavioural auditing. All transcripts are from a single default condition (a standard step-by-step solve prompt; no special system prompt or prefix). This is one of a pair of sets derived from a common transcript pool. Each set contains the same trusted model and one model under investigation; the sets do not disclose how the two investigated models relate… See the full description on the dataset page: https://huggingface.co/datasets/darklord1611/math-eval-transcripts-b.textquestion-answering10K<n<100K0 likes17 downloads2mo agoHugging Face14tonychenxyz /frontier-ai-podcast-transcripts Frontier AI Researcher Podcast Transcripts Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode. Contents 327 episodes 49,286 merged dialogue turns 5,595,982 English tokens using the o200k_base tokenizer 3,942,026 tokens in guest turns Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.tabulartext-generationn<1K0 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.