CoolFace
Datasetpublic

Noothi/huberman-lab-transcripts

Huberman Lab Transcript Dataset Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel. Dataset 438 videos 9,833 transcript chunks ~114 million characters JSONL format Each record contains: ext ideo_id itle Processing The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/huberman-lab-transcripts.

sourceHugging Faceupdated 18d agoView on Hugging Face
0likes77downloads
Dataset Card

Huberman Lab Transcript Dataset

Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel.

Dataset

  • 438 videos
  • 9,833 transcript chunks
  • ~114 million characters
  • JSONL format

Each record contains:

  • ext
  • ideo_id
  • itle

Processing

The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and deduplicating chunks.

Source

The underlying material originates from videos published by the Huberman Lab YouTube channel.

Please independently verify the licensing and redistribution rights applicable to the underlying source material before using this dataset.

Limitations

Caption transcription errors may remain. This dataset should not be considered an authoritative transcript.