datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
massive-yt-edu-queue
Massive YouTube Educational Video Queue
Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours.
Description
This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.massive-yt-edu-transcriptions
Massive YouTube Educational Transcriptions
Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5.
Stats
Videos: 59,355
Characters: 1,539,022,925 (~384M tokens)
Audio hours: 35,890
Model: faster-whisper (CTranslate2) with distil-large-v3.5
Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime
Fields
Field
Description
video_id
YouTube video ID
title
Video title
text
Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.
