CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes9 downloads7mo agoHugging Face02hossam87 /el-mal-el-halal-podcast-subtitles El Mal El Halal Podcast Subtitles Dataset Summary El Mal El Halal Podcast Subtitles is a collection of manual subtitles for 18 episodes of the El Mal El Halal podcast by Eng. Mohamed Aboulnaga, covering Arabic content. This dataset is designed for research on speech processing, translation, semantic search, and Arabic NLP. Total episodes: 18 - untill the date of 03/08/2025 Total segments: 13 970 Total words: 166 505 Total duration: 20 h 50 m 56 s (75 057 s) Average… See the full description on the dataset page: https://huggingface.co/datasets/hossam87/el-mal-el-halal-podcast-subtitles.tabularautomatic-speech-recognition10K<n<100K0 likes8 downloads1y agoHugging Face03tonychenxyz /frontier-ai-podcast-transcripts Frontier AI Researcher Podcast Transcripts Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode. Contents 327 episodes 49,286 merged dialogue turns 5,595,982 English tokens using the o200k_base tokenizer 3,942,026 tokens in guest turns Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.tabulartext-generationn<1K0 likes6 downloads2mo agoHugging Face04hudsongouge /podcast-transcripts-cleaned-phase2gated Podcast Transcripts Cleaned (Phase 2) Updated 2026-07-18T23-19-55Z UTC. Configs Config Rows Description episodes 4,304 Full-episode raw ASR → cleaned transcript (with episode_id / show_id) chunks 5,646 Per-chunk raw → cleaned pairs traces 287,271 Full LM traces (cleaner + evaluator): prompts, reasoning/CoT, tool outputs, pass/fail Cleaned pairs (episodes / chunks) Columns include episode_id, show_id, title, url, instruction… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/podcast-transcripts-cleaned-phase2.tabulartext-generation100K<n<1M1 likes4 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.