CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
0164bits /lex_fridman_podcast_for_llm_vicuna Intro This dataset represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman, is a deep dive into a broad range of topics that touch on science, technology, history, philosophy, and the nature of intelligence, consciousness, love, and power. The guests on the podcast are drawn from a diverse range of fields, providing unique and insightful perspectives on these subjects. The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/64bits/lex_fridman_podcast_for_llm_vicuna.texttext-generation10K<n<100K16 likes146 downloads3y agoHugging Face02RamAnanth1 /lex-fridman-podcasts Dataset Card for Lex Fridman Podcasts Dataset This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model texttext-classificationn<1K6 likes98 downloads4y agoHugging Face03Aditya0619 /lex-fridman-podcast Lex Fridman Podcast Conversations Dataset Dataset Description This dataset contains transcriptions of conversations from the Lex Fridman Podcast, featuring in-depth discussions on artificial intelligence, science, technology, philosophy, and more. The dataset includes 441 transcribed episodes, covering most of the podcast episodes up to January 2025 (excluding 10 episodes). Dataset Structure Features Title: String - The title of the podcast episode… See the full description on the dataset page: https://huggingface.co/datasets/Aditya0619/lex-fridman-podcast.texttext-generationn<1K3 likes36 downloads2y agoHugging Face04BeliefEngines /podcast-transcripts Podcast Transcripts & Belief Graph Structured belief extractions, transcripts, speaker profiles, and embeddings mined from Bitcoin / crypto podcasts by the be-podcast-etl pipeline. Scale (snapshot 2026-04-21) Asset Count Episodes (manifests) 1,551 Podcasts 18 Speakers 876 Persons (enriched profiles) 3,915 Belief shards 66,453 Embeddings (1536-dim) 65,007 Matrices 62,882 Top podcasts: simply-bitcoin (375), the-bitcoin-matrix (264)… See the full description on the dataset page: https://huggingface.co/datasets/BeliefEngines/podcast-transcripts.texttext-generationn<1K0 likes26 downloads5mo agoHugging Face05RamAnanth1 /talkrl-podcast Dataset Card for "talkrl-podcast" This dataset is sourced from the TalkRL Podcast website and contains English transcripts of wonderful TalkRL podcast episodes. The transcripts were generated using OpenAI's base Whisper model texttext-classificationn<1K3 likes22 downloads4y agoHugging Face06ianiket23 /podcast_llama_chat_format Intro This dataset formats an existing podcast dataset (64bits/lex_fridman_podcast_for_llm_vicuna) for llama 3 chat model fine tuning. It represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman. Problems There might be some minor issues during the transcribe phase. Next Step Use whisper to directly load the podcast and transcribe it in this format. texttext-generation10K<n<100K1 likes21 downloads2y agoHugging Face07vanguard-wall /the-vanguard-wall-podcast-episodes The Vanguard Wall Podcast — Episode Catalog Dataset Structured metadata for episodes of The Vanguard Wall Podcast, a long-form interview show with combat veterans, special operators, and first responders. Hosted by US Army veteran Asher Schuler. What this dataset contains Public metadata about each published episode: Episode number, title, description Guest name and credentials (where public) Topic categorization (Combat & Operations, Special Operations, Training &… See the full description on the dataset page: https://huggingface.co/datasets/vanguard-wall/the-vanguard-wall-podcast-episodes.text-classificationn<1K0 likes18 downloads5mo agoHugging Face08rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes9 downloads7mo agoHugging Face09hossam87 /el-mal-el-halal-podcast-subtitles El Mal El Halal Podcast Subtitles Dataset Summary El Mal El Halal Podcast Subtitles is a collection of manual subtitles for 18 episodes of the El Mal El Halal podcast by Eng. Mohamed Aboulnaga, covering Arabic content. This dataset is designed for research on speech processing, translation, semantic search, and Arabic NLP. Total episodes: 18 - untill the date of 03/08/2025 Total segments: 13 970 Total words: 166 505 Total duration: 20 h 50 m 56 s (75 057 s) Average… See the full description on the dataset page: https://huggingface.co/datasets/hossam87/el-mal-el-halal-podcast-subtitles.tabularautomatic-speech-recognition10K<n<100K0 likes8 downloads1y agoHugging Face10tonychenxyz /frontier-ai-podcast-transcripts Frontier AI Researcher Podcast Transcripts Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode. Contents 327 episodes 49,286 merged dialogue turns 5,595,982 English tokens using the o200k_base tokenizer 3,942,026 tokens in guest turns Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.tabulartext-generationn<1K0 likes6 downloads2mo agoHugging Face11ianiket23 /podcast_llama_chat_format-1k Intro This dataset(1K) formats an existing podcast dataset (64bits/lex_fridman_podcast_for_llm_vicuna) for llama 3 chat model fine tuning. It represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman. Problems There might be some minor issues during the transcribe phase. Next Step Use whisper to directly load the podcast and transcribe it in this format. texttext-generation1K<n<10K0 likes5 downloads2y agoHugging Face12hudsongouge /podcast-transcripts-cleaned-phase2gated Podcast Transcripts Cleaned (Phase 2) Updated 2026-07-18T23-19-55Z UTC. Configs Config Rows Description episodes 4,304 Full-episode raw ASR → cleaned transcript (with episode_id / show_id) chunks 5,646 Per-chunk raw → cleaned pairs traces 287,271 Full LM traces (cleaner + evaluator): prompts, reasoning/CoT, tool outputs, pass/fail Cleaned pairs (episodes / chunks) Columns include episode_id, show_id, title, url, instruction… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/podcast-transcripts-cleaned-phase2.tabulartext-generation100K<n<1M1 likes4 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.