CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
0164bits /lex_fridman_podcast_for_llm_vicuna Intro This dataset represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman, is a deep dive into a broad range of topics that touch on science, technology, history, philosophy, and the nature of intelligence, consciousness, love, and power. The guests on the podcast are drawn from a diverse range of fields, providing unique and insightful perspectives on these subjects. The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/64bits/lex_fridman_podcast_for_llm_vicuna.texttext-generation10K<n<100K16 likes146 downloads3y agoHugging Face02YuKuanFu /podcast-dialogue-dataset-shartabular100K<n<1M1 likes106 downloads1y agoHugging Face03RamAnanth1 /lex-fridman-podcasts Dataset Card for Lex Fridman Podcasts Dataset This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model texttext-classificationn<1K6 likes98 downloads4y agoHugging Face04shuyuej /CC-BY-STEMM-Podcast-Transcriptstext10K<n<100K1 likes47 downloads2y agoHugging Face05reapxdev /apple-podcasts-scraper Apple Podcasts Scraper · Shows, Episodes, Genres & Rankings Scrape Apple Podcasts catalog, shows, episodes, top charts, genres, and rankings. HTTP-only iTunes Search API scraper for audio analytics, podcast discovery, and media datasets. Rows in this dataset 2,189 Fields 20 Collector runs behind it 51 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/apple-podcasts-scraper.image1K<n<10K1 likes33 downloads2mo agoHugging Face06filipwx /ted-podcast-finetune LLM Fine-tuning Dataset: TED Talks + Podcasts A structured dataset of transcripts from popular TED Talks and podcasts (Lex Fridman Podcast, Joe Rogan Experience), formatted for LLM fine-tuning. Dataset Summary Property Value Total chunks 2,036 Unique episodes/talks 48 Train split 1,831 records Validation split 204 records Approx. total words 0 Languages English (primary), Portuguese (some TED) Format Chat / Instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/ted-podcast-finetune.text1K<n<10K0 likes21 downloads6mo agoHugging Face07ucalyptus /Ask-ANI-Podcasttext1K<n<10K0 likes17 downloads3y agoHugging Face08shuyuej /CC-BY-STEMM-Podcast-Transcripts-2048text10K<n<100K1 likes16 downloads2y agoHugging Face09instinct-org /espeech_podcasts_chunked_tokenizedgated espeech_podcasts_chunked_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tokenized.tabulartext-to-speech1M<n<10M0 likes15 downloads4mo agoHugging Face10otmanheddouch /andrew-tate-podcasttextn<1K0 likes12 downloads1y agoHugging Face11Gopher-Lab /Sam_Altman_OpenAI_Podcast_XScraper_Example 🔍 X-Twitter Scraper: Real-Time Tweet Search & Scrape Tool Search and scrape X-Twitter for posts by keyword, account, or trending topics.A simple, no-code tool to pull real-time, relevant content in LLM-ready JSON format — perfect for agents, RAG systems, or content workflows. 👉 Start Searching & Scraping on Hugging Face ✨ Features ⚡ Real-Time FetchStream the latest tweets as they’re posted — no delay. 🎯 Flexible SearchSearch by keywords, #hashtags, $cashtags… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/Sam_Altman_OpenAI_Podcast_XScraper_Example.texttext-classificationn<1K0 likes11 downloads1y agoHugging Face12StanKonkin /podcast-assistant-feedbacktextn<1K0 likes8 downloads7mo agoHugging Face13JulianAllen63 /podcastDatatext1K<n<10K0 likes7 downloads3y agoHugging Face14KhangPTT373 /Thuan_podcasttextn<1K0 likes7 downloads1y agoHugging Face15tonychenxyz /frontier-ai-podcast-transcripts Frontier AI Researcher Podcast Transcripts Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode. Contents 327 episodes 49,286 merged dialogue turns 5,595,982 English tokens using the o200k_base tokenizer 3,942,026 tokens in guest turns Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.tabulartext-generationn<1K0 likes6 downloads2mo agoHugging Face16NewEden-Forge /Letterboxd_Podcast_transcripts-sharegpttextn<1K0 likes4 downloads2y agoHugging Face17NewEden-Forge /Numberphile-podcast-sharegpttextn<1K0 likes4 downloads2y agoHugging Face18podcasts-org /podcast_pile_1m_splitgatedtext1M<n<10M0 likes1 downloads11mo agoHugging Face19podcasts-org /podcast_pile_subsets_1m_5m_balancedgatedtext10M<n<100M0 likes1 downloads11mo agoHugging Face20podcasts-org /podcast-pilegatedtabular1K<n<10K0 likes1 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.