datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HLVid
HLVid Dataset
Project Page | Paper | GitHub
HLVid (High-resolution, Long-form Video QA) is a benchmark introduced in the paper "Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing".
It is designed to evaluate Multi-modal Large Language Models (MLLMs) on long-form, high-resolution video understanding. The benchmark features 5-minute videos at 4K resolution, challenging models to handle significant spatiotemporal redundancy while preserving… See the full description on the dataset page: https://huggingface.co/datasets/bfshi/HLVid.HLVOnline_modeldatabase
