datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.5000-podcast-conversations-with-metadata-and-embedding-dataset
🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset
ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications.
This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network.
AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.podcasts_tashkent_dialect_youtube_uzbek_speech_dataset
Tashkent dialect focused podcasts youtube uzbek speech
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.Turkish-Podcast-Merge-v1PODCASTpodcasts_spotifypodcast-audio-slicesTurkish-Podcast-merge-v2PODCAST_maparabic-audio-collection-sudanese-sudan-podcast
Sudan Podcast Arabic Speech Dataset
Dataset Summary
The Sudan Podcast Arabic Speech Dataset is a large-scale Sudanese Arabic speech corpus containing approximately 132 hours of speech recordings and corresponding transcripts, sourced from long-form podcast-style content.
Sudanese Arabic is severely underrepresented in speech technology resources. With over 130 hours of natural, conversational, dialectal speech, this dataset is one of the largest openly available… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-sudan-podcast.PODCASTPODCAST_map_vc_0.5spodcast-tokenized-bg3.5-enj5lex_fridman_podcast_for_llm_vicuna
Intro
This dataset represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman, is a deep dive into a broad range of topics that touch on science, technology, history, philosophy, and the nature of intelligence, consciousness, love, and power. The guests on the podcast are drawn from a diverse range of fields, providing unique and insightful perspectives on these subjects.
The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/64bits/lex_fridman_podcast_for_llm_vicuna.podcast-tokenized-bg2.5-enj4.5arabic-audio-collection-syrian-podcast
Syrian Postcast Arabic Speech Dataset
Dataset Summary
The Syrian Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 116 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-syrian-podcast.mellow-podcast-datalex-fridman-podcast
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast.podcast_fillers_by_license
Some Podcasts
Podcasts are taken from the PodcastFillers dataset. The PodcastFillers dataset consists of 199 full-length podcast episodes in English with manually annotated filler words and automatically generated transcripts. The podcast audio recordings, sourced from SoundCloud, are CC-licensed, gender-balanced, and total 145 hours of audio from over 350 speakers.
[!TIP]
This dataset doesn't upload the PodcastFillers annotations, which are under a non-commercial license. See here… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/podcast_fillers_by_license.ESpeech-podcasts
Podcasts Audio Dataset
Dataset Description
This dataset contains 3200 hours of processed audio segments extracted from various podcasts with corresponding metadata. Each audio file represents a segment from podcast episodes, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Total Duration: 3200 hours of speech
Task: TTS, ASR, Quality Assessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-podcasts.podcast-dialogue-dataset-sharlex-fridman-podcasts
Dataset Card for Lex Fridman Podcasts Dataset
This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model
lex_fridman_podcast
Dataset Card for "lex_fridman_podcast"
Dataset Summary
This dataset contains transcripts from the Lex Fridman podcast (Episodes 1 to 325).
The transcripts were generated using OpenAI Whisper (large model) and made publicly available at: https://karpathy.ai/lexicap/index.html.
Languages
English
Dataset Structure
The dataset contains around 803K entries, consisting of audio transcripts generated from episodes 1 to 325 of the Lex Fridman… See the full description on the dataset page: https://huggingface.co/datasets/nmac/lex_fridman_podcast.PODCAST_map_vc_0.1stech-podcast-audio-30s
Tech Podcast Audio 30s
This is a Japanese speech corpus derived from technology-related podcast audio distributed through publicly accessible podcast RSS feeds. Podcast episodes were segmented into clips of up to 29.9 seconds using voice activity detection.
The combined dataset contains 1,129,703 accepted clips totaling 5,516.14 hours (approximately 329 GiB of Parquet files). Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/tech-podcast-audio-30s.tamil-english-podcast-diarization
Tamil-English Code-Mixed Podcast Diarization Dataset
Dataset Summary
This dataset contains long-form Tamil-English code-mixed podcast recordings
annotated for speaker diarization research. The recordings consist of natural
conversational speech with multiple speakers and realistic acoustic conditions,
making the dataset suitable for evaluating diarization pipelines in
real-world scenarios.
The dataset is intended to support research in:
Speaker diarization
Code-mixed… See the full description on the dataset page: https://huggingface.co/datasets/Rangasuthan/tamil-english-podcast-diarization.PodcastCleanDatapodcast-dataClean-Podcast
Clean Podcast & Movie Teacher Subsets
Clean 48 kHz / 16-bit / mono speech chunks used as teacher targets for the call-centre speech-restoration finetune
(Malaysian/Singaporean telephony domain). These are the clean references only — the model
learns to reconstruct this clean speech from a telephony-degraded version that is synthesised
on the fly at training time (band-limiting, GSM/G.711-µ-law/MP3 codecs, line noise, VoIP
dropouts). No degraded audio is stored here.
Each subset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Clean-Podcast.arabic-audio-collection-libyan-a7rar-podcast
Libyan Postcast Arabic Speech Dataset
Dataset Summary
The Libyan Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 26 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-libyan-a7rar-podcast.
