datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qvhighlights-1fps
QVHighlights 1fps — Preprocessed Frames
Preprocessed version of the QVHighlights dataset for temporal video grounding.
Videos are extracted at 1fps, resized to 384×384 JPEG, ready for training without any video I/O at runtime.
Contents
File
Description
annotations_train.jsonl
7445 train annotations
annotations_val.jsonl
1550 val annotations
frames folder
Train frames batch 0000–1000
frames_000000_001000.tar
Train frames batch 0000–1000… See the full description on the dataset page: https://huggingface.co/datasets/shaunmarvell/qvhighlights-1fps.qvhighlights-valqvhighlights-testr1-qvhighlights-8619qvhighlights-25frames-testqvhighlights-yesno-qa
