datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chapter-llama
VidChapters Dataset for Chapter-Llama
This repository contains the dataset used in the paper "Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs" (CVPR 2025).
Overview
VidChapters-7M is a large-scale dataset for video chaptering, containing:
817k videos with ASR data (20GB)
Captions extracted from videos using various sampling strategies
Chapter annotations with timestamps and titles
Data Structure
The dataset is organized as follows:
ASR… See the full description on the dataset page: https://huggingface.co/datasets/lucas-ventura/chapter-llama.Activitynet_Video_llamaLlama4_RLHFLlama4_SFT
