datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eluniversoraro-podcastlex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.5000-podcast-conversations-with-metadata-and-embedding-dataset
🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset
ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications.
This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network.
AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.podcast-tokenized-bg3.5-enj5-with-speaker-embeddings
podcast-tokenized-bg3.5-enj5-with-speaker-embeddings
This dataset extends TTS-AGI/podcast-tokenized-bg3.5-enj5 with speaker embeddings, cosine similarity scores, speaker cluster assignments, and reference-match flags for each sample.
What was added
Each sample's JSON metadata is augmented with the following fields:
Field
Type
Description
target_speaker_embedding
list[float] (128-dim)
L2-normalized speaker embedding of the target audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/podcast-tokenized-bg3.5-enj5-with-speaker-embeddings.podcastvideosdotgov-podcast-samplepodcasts_tashkent_dialect_youtube_uzbek_speech_dataset
Tashkent dialect focused podcasts youtube uzbek speech
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.Turkish-Podcast-Merge-v1PODCASTmalaysian-podcast-youtube
Crawl Youtube Malaysian Podcast
With total 19092 audio files, total 2233.8 hours.
how to download
huggingface-cli download --repo-type dataset \
--include '*.z*' \
--local-dir './' \
malaysia-ai/malaysian-podcast-youtube
wget https://www.7-zip.org/a/7z2301-linux-x64.tar.xz
tar -xf 7z2301-linux-x64.tar.xz
~/7zz x malaysian-podcast.zip -y -mmt40
Licensing
All the videos, songs, images, and graphics used in the video belong to their respective owners and I does… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-podcast-youtube.podcasts_spotifyTurkish-Podcast-merge-v2espeech_podcasts_chunked_tts_train
espeech_podcasts_chunked_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tts_train.podcast-audio-sliceshindi_podcast_datasingaporean-podcast-youtube
Crawl Youtube Singaporean Podcast
With total 3451 audio files, total 1254.6 hours.
how to download
huggingface-cli download --repo-type dataset \
--include '*.z*' \
--local-dir './' \
malaysia-ai/singaporean-podcast-youtube
wget https://www.7-zip.org/a/7z2301-linux-x64.tar.xz
tar -xf 7z2301-linux-x64.tar.xz
~/7zz x sg-podcast.zip -y -mmt40
Licensing
All the videos, songs, images, and graphics used in the video belong to their respective owners and I does not… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/singaporean-podcast-youtube.PODCAST_maparabic-audio-collection-sudanese-sudan-podcast
Sudan Podcast Arabic Speech Dataset
Dataset Summary
The Sudan Podcast Arabic Speech Dataset is a large-scale Sudanese Arabic speech corpus containing approximately 132 hours of speech recordings and corresponding transcripts, sourced from long-form podcast-style content.
Sudanese Arabic is severely underrepresented in speech technology resources. With over 130 hours of natural, conversational, dialectal speech, this dataset is one of the largest openly available… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-sudan-podcast.podcast-tokenized-bg3.5-enj5PODCASTpreprocessed_spanish_podcastsPODCAST_map_vc_0.5slex_fridman_podcast_for_llm_vicuna
Intro
This dataset represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman, is a deep dive into a broad range of topics that touch on science, technology, history, philosophy, and the nature of intelligence, consciousness, love, and power. The guests on the podcast are drawn from a diverse range of fields, providing unique and insightful perspectives on these subjects.
The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/64bits/lex_fridman_podcast_for_llm_vicuna.podcast-tokenized-bg2.5-enj4.5arabic-audio-collection-syrian-podcast
Syrian Postcast Arabic Speech Dataset
Dataset Summary
The Syrian Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 116 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-syrian-podcast.MSP_podcastVerbalVerse_of_Podcastmellow-podcast-datalex-fridman-podcast
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast.edu_podcast_news_mp3_files
