datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lex_fridman_podcast_for_llm_vicuna
Intro
This dataset represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman, is a deep dive into a broad range of topics that touch on science, technology, history, philosophy, and the nature of intelligence, consciousness, love, and power. The guests on the podcast are drawn from a diverse range of fields, providing unique and insightful perspectives on these subjects.
The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/64bits/lex_fridman_podcast_for_llm_vicuna.lex-fridman-podcasts
Dataset Card for Lex Fridman Podcasts Dataset
This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model
lex-fridman-podcast
Lex Fridman Podcast Conversations Dataset
Dataset Description
This dataset contains transcriptions of conversations from the Lex Fridman Podcast, featuring in-depth discussions on artificial intelligence, science, technology, philosophy, and more. The dataset includes 441 transcribed episodes, covering most of the podcast episodes up to January 2025 (excluding 10 episodes).
Dataset Structure
Features
Title: String - The title of the podcast episode… See the full description on the dataset page: https://huggingface.co/datasets/Aditya0619/lex-fridman-podcast.podcast-transcripts
Podcast Transcripts & Belief Graph
Structured belief extractions, transcripts, speaker profiles, and embeddings
mined from Bitcoin / crypto podcasts by the
be-podcast-etl pipeline.
Scale (snapshot 2026-04-21)
Asset
Count
Episodes (manifests)
1,551
Podcasts
18
Speakers
876
Persons (enriched profiles)
3,915
Belief shards
66,453
Embeddings (1536-dim)
65,007
Matrices
62,882
Top podcasts: simply-bitcoin (375), the-bitcoin-matrix (264)… See the full description on the dataset page: https://huggingface.co/datasets/BeliefEngines/podcast-transcripts.talkrl-podcast
Dataset Card for "talkrl-podcast"
This dataset is sourced from the TalkRL Podcast website and contains English transcripts of wonderful TalkRL podcast episodes. The transcripts were generated using OpenAI's base Whisper model
podcast_llama_chat_format
Intro
This dataset formats an existing podcast dataset (64bits/lex_fridman_podcast_for_llm_vicuna) for llama 3 chat model fine tuning.
It represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman.
Problems
There might be some minor issues during the transcribe phase.
Next Step
Use whisper to directly load the podcast and transcribe it in this format.
the-vanguard-wall-podcast-episodes
The Vanguard Wall Podcast — Episode Catalog Dataset
Structured metadata for episodes of The Vanguard Wall Podcast, a long-form interview show with combat veterans, special operators, and first responders. Hosted by US Army veteran Asher Schuler.
What this dataset contains
Public metadata about each published episode:
Episode number, title, description
Guest name and credentials (where public)
Topic categorization (Combat & Operations, Special Operations, Training &… See the full description on the dataset page: https://huggingface.co/datasets/vanguard-wall/the-vanguard-wall-podcast-episodes.podcast-transcripts
Podcast Transcripts Dataset
This dataset contains transcripts from Bitcoin and cryptocurrency podcasts,
processed by the belief-engines ETL pipeline.
Files
transcripts.parquet - Full episode transcripts with speaker diarization
transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings
Schema
transcripts.parquet
episode_id: Unique episode identifier
podcast_slug: Podcast name slug
episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.el-mal-el-halal-podcast-subtitles
El Mal El Halal Podcast Subtitles
Dataset Summary
El Mal El Halal Podcast Subtitles is a collection of manual subtitles for 18 episodes of the El Mal El Halal podcast by Eng. Mohamed Aboulnaga, covering Arabic content. This dataset is designed for research on speech processing, translation, semantic search, and Arabic NLP.
Total episodes: 18 - untill the date of 03/08/2025
Total segments: 13 970
Total words: 166 505
Total duration: 20 h 50 m 56 s (75 057 s)
Average… See the full description on the dataset page: https://huggingface.co/datasets/hossam87/el-mal-el-halal-podcast-subtitles.frontier-ai-podcast-transcripts
Frontier AI Researcher Podcast Transcripts
Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode.
Contents
327 episodes
49,286 merged dialogue turns
5,595,982 English tokens using the o200k_base tokenizer
3,942,026 tokens in guest turns
Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.podcast_llama_chat_format-1k
Intro
This dataset(1K) formats an existing podcast dataset (64bits/lex_fridman_podcast_for_llm_vicuna) for llama 3 chat model fine tuning.
It represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman.
Problems
There might be some minor issues during the transcribe phase.
Next Step
Use whisper to directly load the podcast and transcribe it in this format.
podcast-transcripts-cleaned-phase2
Podcast Transcripts Cleaned (Phase 2)
Updated 2026-07-18T23-19-55Z UTC.
Configs
Config
Rows
Description
episodes
4,304
Full-episode raw ASR → cleaned transcript (with episode_id / show_id)
chunks
5,646
Per-chunk raw → cleaned pairs
traces
287,271
Full LM traces (cleaner + evaluator): prompts, reasoning/CoT, tool outputs, pass/fail
Cleaned pairs (episodes / chunks)
Columns include episode_id, show_id, title, url, instruction… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/podcast-transcripts-cleaned-phase2.
