datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
el-mal-el-halal-podcast-subtitles
El Mal El Halal Podcast Subtitles
Dataset Summary
El Mal El Halal Podcast Subtitles is a collection of manual subtitles for 18 episodes of the El Mal El Halal podcast by Eng. Mohamed Aboulnaga, covering Arabic content. This dataset is designed for research on speech processing, translation, semantic search, and Arabic NLP.
Total episodes: 18 - untill the date of 03/08/2025
Total segments: 13 970
Total words: 166 505
Total duration: 20 h 50 m 56 s (75 057 s)
Average… See the full description on the dataset page: https://huggingface.co/datasets/hossam87/el-mal-el-halal-podcast-subtitles.frontier-ai-podcast-transcripts
Frontier AI Researcher Podcast Transcripts
Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode.
Contents
327 episodes
49,286 merged dialogue turns
5,595,982 English tokens using the o200k_base tokenizer
3,942,026 tokens in guest turns
Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.
