Noothi/huberman-lab-transcripts
Huberman Lab Transcript Dataset Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel. Dataset 438 videos 9,833 transcript chunks ~114 million characters JSONL format Each record contains: ext ideo_id itle Processing The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/huberman-lab-transcripts.
083
1---2pretty_name: Huberman Lab Transcript Dataset3language:4 - en5task_categories:6 - text-generation7tags:8 - transcripts9 - youtube10 - huberman-lab11 - llm12size_categories:13 - 10M<n<100M14---15 16# Huberman Lab Transcript Dataset17 18Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel.19 20## Dataset21 22- 438 videos23- 9,833 transcript chunks24- ~114 million characters25- JSONL format26 27Each record contains:28 29- ext30- ideo_id31- itle32 33## Processing34 35The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and deduplicating chunks.36 37## Source38 39The underlying material originates from videos published by the Huberman Lab YouTube channel.40 41Please independently verify the licensing and redistribution rights applicable to the underlying source material before using this dataset.42 43## Limitations44 45Caption transcription errors may remain. This dataset should not be considered an authoritative transcript.46
47 