CoolFace
Datasetpublic

Whispering-GPT/lex-fridman-podcast

Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id:… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast.

sourceHugging Faceupdated 3y agoView on Hugging Face
12likes125downloads
Dataset Card

Dataset Card for "lexFridmanPodcast-transcript-audio"

Table of Contents

Dataset Description

Dataset Summary

This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.

Languages

  • —Language: English

Dataset Structure

The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.

Data Fields

The dataset is composed by:

  • —id: Id of the youtube video.
  • —channel: Name of the channel.
  • —channel\_id: Id of the youtube channel.
  • —title: Title given to the video.
  • —categories: Category of the video.
  • —description: Description added by the author.
  • —text: Whole transcript of the video.
  • —segments: A list with the time and transcription of the video.
  • —start: When started the trancription.
  • —end: When the transcription ends.
  • —text: The text of the transcription.

Data Splits

  • —Train split.

Dataset Creation

Source Data

The transcriptions are from the videos of Lex Fridman Podcast

Contributions

Thanks to Whispering-GPT organization for adding this dataset.