datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
podcast-dialogue-dataset-sharapple-podcasts-scraper
Apple Podcasts Scraper · Shows, Episodes, Genres & Rankings
Scrape Apple Podcasts catalog, shows, episodes, top charts, genres, and rankings. HTTP-only iTunes Search API scraper for audio analytics, podcast discovery, and media datasets.
Rows in this dataset
2,189
Fields
20
Collector runs behind it
51
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/apple-podcasts-scraper.espeech_podcasts_chunked_tokenized
espeech_podcasts_chunked_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tokenized.frontier-ai-podcast-transcripts
Frontier AI Researcher Podcast Transcripts
Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode.
Contents
327 episodes
49,286 merged dialogue turns
5,595,982 English tokens using the o200k_base tokenizer
3,942,026 tokens in guest turns
Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.podcast-pile
