datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lex_fridman_podcast_for_llm_vicuna
Intro
This dataset represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman, is a deep dive into a broad range of topics that touch on science, technology, history, philosophy, and the nature of intelligence, consciousness, love, and power. The guests on the podcast are drawn from a diverse range of fields, providing unique and insightful perspectives on these subjects.
The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/64bits/lex_fridman_podcast_for_llm_vicuna.podcast-dialogue-dataset-sharlex-fridman-podcasts
Dataset Card for Lex Fridman Podcasts Dataset
This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model
CC-BY-STEMM-Podcast-Transcriptsapple-podcasts-scraper
Apple Podcasts Scraper · Shows, Episodes, Genres & Rankings
Scrape Apple Podcasts catalog, shows, episodes, top charts, genres, and rankings. HTTP-only iTunes Search API scraper for audio analytics, podcast discovery, and media datasets.
Rows in this dataset
2,189
Fields
20
Collector runs behind it
51
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/apple-podcasts-scraper.ted-podcast-finetune
LLM Fine-tuning Dataset: TED Talks + Podcasts
A structured dataset of transcripts from popular TED Talks and podcasts (Lex Fridman Podcast, Joe Rogan Experience), formatted for LLM fine-tuning.
Dataset Summary
Property
Value
Total chunks
2,036
Unique episodes/talks
48
Train split
1,831 records
Validation split
204 records
Approx. total words
0
Languages
English (primary), Portuguese (some TED)
Format
Chat / Instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/ted-podcast-finetune.Ask-ANI-PodcastCC-BY-STEMM-Podcast-Transcripts-2048espeech_podcasts_chunked_tokenized
espeech_podcasts_chunked_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tokenized.andrew-tate-podcastSam_Altman_OpenAI_Podcast_XScraper_Example
🔍 X-Twitter Scraper: Real-Time Tweet Search & Scrape Tool
Search and scrape X-Twitter for posts by keyword, account, or trending topics.A simple, no-code tool to pull real-time, relevant content in LLM-ready JSON format — perfect for agents, RAG systems, or content workflows.
👉 Start Searching & Scraping on Hugging Face
✨ Features
⚡ Real-Time FetchStream the latest tweets as they’re posted — no delay.
🎯 Flexible SearchSearch by keywords, #hashtags, $cashtags… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/Sam_Altman_OpenAI_Podcast_XScraper_Example.podcast-assistant-feedbackpodcastDataThuan_podcastfrontier-ai-podcast-transcripts
Frontier AI Researcher Podcast Transcripts
Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode.
Contents
327 episodes
49,286 merged dialogue turns
5,595,982 English tokens using the o200k_base tokenizer
3,942,026 tokens in guest turns
Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.Letterboxd_Podcast_transcripts-sharegptNumberphile-podcast-sharegptpodcast_pile_1m_splitpodcast_pile_subsets_1m_5m_balancedpodcast-pile
