CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Whispering-GPT /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes2.1k downloads4y agoHugging Face02ReadyAi /5000-podcast-conversations-with-metadata-and-embedding-dataset 🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications. This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network. AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.text10K<n<100K8 likes1.2k downloads1y agoHugging Face03islomov /podcasts_tashkent_dialect_youtube_uzbek_speech_dataset Tashkent dialect focused podcasts youtube uzbek speech Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K6 likes389 downloads1y agoHugging Face04Aynursusuz /Turkish-Podcast-Merge-v1audio100K<n<1M2 likes372 downloads7mo agoHugging Face05AbstractTTS /PODCASTaudio100K<n<1M20 likes349 downloads2y agoHugging Face06laudite-ufg /podcasts_spotifyaudio100K<n<1M0 likes291 downloads1y agoHugging Face07Adam429 /podcast-audio-slicestext0 likes243 downloads6mo agoHugging Face08Aynursusuz /Turkish-Podcast-merge-v2audio100K<n<1M1 likes213 downloads9mo agoHugging Face09windcrossroad /PODCAST_mapaudio10K<n<100K0 likes211 downloads2y agoHugging Face10oddadmix /arabic-audio-collection-sudanese-sudan-podcast Sudan Podcast Arabic Speech Dataset Dataset Summary The Sudan Podcast Arabic Speech Dataset is a large-scale Sudanese Arabic speech corpus containing approximately 132 hours of speech recordings and corresponding transcripts, sourced from long-form podcast-style content. Sudanese Arabic is severely underrepresented in speech technology resources. With over 130 hours of natural, conversational, dialectal speech, this dataset is one of the largest openly available… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-sudan-podcast.audiotext-to-speech10K<n<100K1 likes196 downloads2mo agoHugging Face11windcrossroad /PODCASTaudio10K<n<100K0 likes174 downloads2y agoHugging Face12AdoCleanCode /PODCAST_map_vc_0.5stext10K<n<100K0 likes161 downloads7mo agoHugging Face13TTS-AGI /podcast-tokenized-bg3.5-enj5text1M<n<10M0 likes146 downloads7mo agoHugging Face1464bits /lex_fridman_podcast_for_llm_vicuna Intro This dataset represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman, is a deep dive into a broad range of topics that touch on science, technology, history, philosophy, and the nature of intelligence, consciousness, love, and power. The guests on the podcast are drawn from a diverse range of fields, providing unique and insightful perspectives on these subjects. The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/64bits/lex_fridman_podcast_for_llm_vicuna.texttext-generation10K<n<100K16 likes145 downloads3y agoHugging Face15TTS-AGI /podcast-tokenized-bg2.5-enj4.5text10M<n<100M1 likes143 downloads7mo agoHugging Face16oddadmix /arabic-audio-collection-syrian-podcast Syrian Postcast Arabic Speech Dataset Dataset Summary The Syrian Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 116 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-syrian-podcast.audiotext-to-speech10K<n<100K2 likes142 downloads3mo agoHugging Face17anonymousforemotion /mellow-podcast-dataaudio1K<n<10K0 likes127 downloads1y agoHugging Face18Whispering-GPT /lex-fridman-podcast Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast.textautomatic-speech-recognitionn<1K12 likes117 downloads3y agoHugging Face19ylacombe /podcast_fillers_by_license Some Podcasts Podcasts are taken from the PodcastFillers dataset. The PodcastFillers dataset consists of 199 full-length podcast episodes in English with manually annotated filler words and automatically generated transcripts. The podcast audio recordings, sourced from SoundCloud, are CC-licensed, gender-balanced, and total 145 hours of audio from over 350 speakers. [!TIP] This dataset doesn't upload the PodcastFillers annotations, which are under a non-commercial license. See here… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/podcast_fillers_by_license.audion<1K0 likes114 downloads2y agoHugging Face20ESpeech /ESpeech-podcastsgated Podcasts Audio Dataset Dataset Description This dataset contains 3200 hours of processed audio segments extracted from various podcasts with corresponding metadata. Each audio file represents a segment from podcast episodes, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Total Duration: 3200 hours of speech Task: TTS, ASR, Quality Assessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-podcasts.audiotext-to-speech10M<n<100M14 likes113 downloads10mo agoHugging Face21YuKuanFu /podcast-dialogue-dataset-shartabular100K<n<1M1 likes106 downloads1y agoHugging Face22RamAnanth1 /lex-fridman-podcasts Dataset Card for Lex Fridman Podcasts Dataset This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model texttext-classificationn<1K6 likes100 downloads4y agoHugging Face23nmac /lex_fridman_podcast Dataset Card for "lex_fridman_podcast" Dataset Summary This dataset contains transcripts from the Lex Fridman podcast (Episodes 1 to 325). The transcripts were generated using OpenAI Whisper (large model) and made publicly available at: https://karpathy.ai/lexicap/index.html. Languages English Dataset Structure The dataset contains around 803K entries, consisting of audio transcripts generated from episodes 1 to 325 of the Lex Fridman… See the full description on the dataset page: https://huggingface.co/datasets/nmac/lex_fridman_podcast.textautomatic-speech-recognition100K<n<1M9 likes96 downloads4y agoHugging Face24AdoCleanCode /PODCAST_map_vc_0.1stext10K<n<100K0 likes92 downloads7mo agoHugging Face25KeisukeMiyamoto /tech-podcast-audio-30sgated Tech Podcast Audio 30s This is a Japanese speech corpus derived from technology-related podcast audio distributed through publicly accessible podcast RSS feeds. Podcast episodes were segmented into clips of up to 29.9 seconds using voice activity detection. The combined dataset contains 1,129,703 accepted clips totaling 5,516.14 hours (approximately 329 GiB of Parquet files). Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/tech-podcast-audio-30s.audioautomatic-speech-recognition1M<n<10M0 likes89 downloads5d agoHugging Face26Rangasuthan /tamil-english-podcast-diarization Tamil-English Code-Mixed Podcast Diarization Dataset Dataset Summary This dataset contains long-form Tamil-English code-mixed podcast recordings annotated for speaker diarization research. The recordings consist of natural conversational speech with multiple speakers and realistic acoustic conditions, making the dataset suitable for evaluating diarization pipelines in real-world scenarios. The dataset is intended to support research in: Speaker diarization Code-mixed… See the full description on the dataset page: https://huggingface.co/datasets/Rangasuthan/tamil-english-podcast-diarization.audioautomatic-speech-recognitionn<1K1 likes85 downloads7mo agoHugging Face27Mohsen21 /PodcastCleanDatatabular1K<n<10K0 likes82 downloads2y agoHugging Face28marcoyang /podcast-datatext0 likes74 downloads11mo agoHugging Face29Scicom-intl /Clean-Podcast Clean Podcast & Movie Teacher Subsets Clean 48 kHz / 16-bit / mono speech chunks used as teacher targets for the call-centre speech-restoration finetune (Malaysian/Singaporean telephony domain). These are the clean references only — the model learns to reconstruct this clean speech from a telephony-degraded version that is synthesised on the fly at training time (band-limiting, GSM/G.711-µ-law/MP3 codecs, line noise, VoIP dropouts). No degraded audio is stored here. Each subset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Clean-Podcast.tabularaudio-to-audio10K<n<100K0 likes66 downloads2mo agoHugging Face30oddadmix /arabic-audio-collection-libyan-a7rar-podcast Libyan Postcast Arabic Speech Dataset Dataset Summary The Libyan Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 26 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-libyan-a7rar-podcast.audiotext-to-speech1K<n<10K1 likes65 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.