bfshi/HLVid
HLVid Dataset Project Page | Paper | GitHub HLVid (High-resolution, Long-form Video QA) is a benchmark introduced in the paper "Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing". It is designed to evaluate Multi-modal Large Language Models (MLLMs) on long-form, high-resolution video understanding. The benchmark features 5-minute videos at 4K resolution, challenging models to handle significant spatiotemporal redundancy while… See the full description on the dataset page: https://huggingface.co/datasets/bfshi/HLVid.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face