datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube_subs_howto100M
Dataset Card for youtube_subs_howto100M
Dataset Summary
The youtube_subs_howto100M dataset is an English-language dataset of instruction-response pairs extracted from 309136 YouTube videos. The dataset was orignally inspired by and sourced from the HowTo100M dataset, which was developed for natural language search for video clips.
Supported Tasks and Leaderboards
conversational: The dataset can be used to train a model for instruction(request) and a long form… See the full description on the dataset page: https://huggingface.co/datasets/totuta/youtube_subs_howto100M.howto100m_captions_with_verb_nounshowto100m
HowTo100M
105,128 videos (11721.4 GB) with 104,584 subtitle files, downloaded at source
quality and re-hosted for direct use — no more dead YouTube links, no more flaky
downloader scripts.
Coverage: 105,128 of the 1,238,911 source video IDs (8.5%). 19,160 source videos were unavailable at fetch time (private, removed, members-only or geo-blocked) and are excluded. The dataset is refreshed as more videos are delivered.
What's inside
metadata.jsonl — one row per… See the full description on the dataset page: https://huggingface.co/datasets/TornadoLabs/howto100m.HowTo100M-subtitles-small
HowTo100M-subtitles-small
The subtitles from a subset of the HowTo100M dataset.
