datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HLVid
HLVid Dataset
Project Page | Paper | GitHub
HLVid (High-resolution, Long-form Video QA) is a benchmark introduced in the paper "Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing".
It is designed to evaluate Multi-modal Large Language Models (MLLMs) on long-form, high-resolution video understanding. The benchmark features 5-minute videos at 4K resolution, challenging models to handle significant spatiotemporal redundancy while preserving… See the full description on the dataset page: https://huggingface.co/datasets/bfshi/HLVid.HLV-1K
🎬 HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
📖 Introduction
HLV-1K is a comprehensive benchmark designed to evaluate the capabilities of multimodal large language models (MLLMs) in understanding hour-long videos with time-specific queries. Unlike existing video understanding benchmarks that focus on short clips, HLV-1K addresses the critical challenge of long-term video comprehension by providing:
🕐 Hour-long… See the full description on the dataset page: https://huggingface.co/datasets/ZouHQ/HLV-1K.hlv_onlinetwitter-TianxinKitten-2025.03.24-1904021703262417406-hlV5u4Tp9dTcocab-part2twitter-TianxinKitten-2025.03.24-1904021703262417406-hlV5u4Tp9dTcocab-part3HLVOnline_modeldatabaseHLVPKeaWtwitter-TianxinKitten-2025.03.24-1904021703262417406-hlV5u4Tp9dTcocab-part1
