CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Juancarlosy /eluniversoraro-podcastaudion<1K0 likes4.3k downloads3mo agoHugging Face02Whispering-GPT /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes2.1k downloads4y agoHugging Face03ReadyAi /5000-podcast-conversations-with-metadata-and-embedding-dataset 🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications. This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network. AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.text10K<n<100K8 likes1.2k downloads1y agoHugging Face04TTS-AGI /podcast-tokenized-bg3.5-enj5-with-speaker-embeddings podcast-tokenized-bg3.5-enj5-with-speaker-embeddings This dataset extends TTS-AGI/podcast-tokenized-bg3.5-enj5 with speaker embeddings, cosine similarity scores, speaker cluster assignments, and reference-match flags for each sample. What was added Each sample's JSON metadata is augmented with the following fields: Field Type Description target_speaker_embedding list[float] (128-dim) L2-normalized speaker embedding of the target audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/podcast-tokenized-bg3.5-enj5-with-speaker-embeddings.text-to-speech0 likes1.1k downloads3mo agoHugging Face05EQ4You /podcastvideos4 likes915 downloads2y agoHugging Face06alea-institute /dotgov-podcast-sampleaudion<1K0 likes418 downloads2y agoHugging Face07islomov /podcasts_tashkent_dialect_youtube_uzbek_speech_dataset Tashkent dialect focused podcasts youtube uzbek speech Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K6 likes394 downloads1y agoHugging Face08Aynursusuz /Turkish-Podcast-Merge-v1audio100K<n<1M2 likes371 downloads7mo agoHugging Face09AbstractTTS /PODCASTaudio100K<n<1M20 likes333 downloads2y agoHugging Face10malaysia-ai /malaysian-podcast-youtube Crawl Youtube Malaysian Podcast With total 19092 audio files, total 2233.8 hours. how to download huggingface-cli download --repo-type dataset \ --include '*.z*' \ --local-dir './' \ malaysia-ai/malaysian-podcast-youtube wget https://www.7-zip.org/a/7z2301-linux-x64.tar.xz tar -xf 7z2301-linux-x64.tar.xz ~/7zz x malaysian-podcast.zip -y -mmt40 Licensing All the videos, songs, images, and graphics used in the video belong to their respective owners and I does… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-podcast-youtube.2 likes311 downloads1y agoHugging Face11laudite-ufg /podcasts_spotifyaudio100K<n<1M0 likes290 downloads1y agoHugging Face12Aynursusuz /Turkish-Podcast-merge-v2audio100K<n<1M1 likes284 downloads9mo agoHugging Face13instinct-org /espeech_podcasts_chunked_tts_traingated espeech_podcasts_chunked_tts_train This is a gated Russian TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tts_train.audiotext-to-speech0 likes253 downloads4mo agoHugging Face14Adam429 /podcast-audio-slicestext0 likes236 downloads6mo agoHugging Face15CanopyLabsEliasF /hindi_podcast_data0 likes228 downloads1y agoHugging Face16malaysia-ai /singaporean-podcast-youtube Crawl Youtube Singaporean Podcast With total 3451 audio files, total 1254.6 hours. how to download huggingface-cli download --repo-type dataset \ --include '*.z*' \ --local-dir './' \ malaysia-ai/singaporean-podcast-youtube wget https://www.7-zip.org/a/7z2301-linux-x64.tar.xz tar -xf 7z2301-linux-x64.tar.xz ~/7zz x sg-podcast.zip -y -mmt40 Licensing All the videos, songs, images, and graphics used in the video belong to their respective owners and I does not… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/singaporean-podcast-youtube.0 likes213 downloads1y agoHugging Face17windcrossroad /PODCAST_mapaudio10K<n<100K0 likes207 downloads2y agoHugging Face18oddadmix /arabic-audio-collection-sudanese-sudan-podcast Sudan Podcast Arabic Speech Dataset Dataset Summary The Sudan Podcast Arabic Speech Dataset is a large-scale Sudanese Arabic speech corpus containing approximately 132 hours of speech recordings and corresponding transcripts, sourced from long-form podcast-style content. Sudanese Arabic is severely underrepresented in speech technology resources. With over 130 hours of natural, conversational, dialectal speech, this dataset is one of the largest openly available… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-sudan-podcast.audiotext-to-speech10K<n<100K1 likes197 downloads2mo agoHugging Face19TTS-AGI /podcast-tokenized-bg3.5-enj5text1M<n<10M0 likes190 downloads7mo agoHugging Face20windcrossroad /PODCASTaudio10K<n<100K0 likes174 downloads2y agoHugging Face21javi22 /preprocessed_spanish_podcasts0 likes168 downloads8mo agoHugging Face22AdoCleanCode /PODCAST_map_vc_0.5stext10K<n<100K0 likes158 downloads7mo agoHugging Face2364bits /lex_fridman_podcast_for_llm_vicuna Intro This dataset represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman, is a deep dive into a broad range of topics that touch on science, technology, history, philosophy, and the nature of intelligence, consciousness, love, and power. The guests on the podcast are drawn from a diverse range of fields, providing unique and insightful perspectives on these subjects. The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/64bits/lex_fridman_podcast_for_llm_vicuna.texttext-generation10K<n<100K16 likes144 downloads3y agoHugging Face24TTS-AGI /podcast-tokenized-bg2.5-enj4.5text10M<n<100M1 likes143 downloads7mo agoHugging Face25oddadmix /arabic-audio-collection-syrian-podcast Syrian Postcast Arabic Speech Dataset Dataset Summary The Syrian Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 116 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-syrian-podcast.audiotext-to-speech10K<n<100K2 likes140 downloads3mo agoHugging Face26CLAPv2 /MSP_podcastimage100K<n<1M1 likes130 downloads1y agoHugging Face27VerbalVerseOfPodcast /VerbalVerse_of_Podcast0 likes129 downloads1y agoHugging Face28anonymousforemotion /mellow-podcast-dataaudio1K<n<10K0 likes127 downloads1y agoHugging Face29Whispering-GPT /lex-fridman-podcast Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast.textautomatic-speech-recognitionn<1K12 likes123 downloads3y agoHugging Face30EQ4You /edu_podcast_news_mp3_files0 likes121 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.