CoolFace
Datasetpublic

bfshi/HLVid

HLVid Dataset Project Page | Paper | GitHub HLVid (High-resolution, Long-form Video QA) is a benchmark introduced in the paper "Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing". It is designed to evaluate Multi-modal Large Language Models (MLLMs) on long-form, high-resolution video understanding. The benchmark features 5-minute videos at 4K resolution, challenging models to handle significant spatiotemporal redundancy while… See the full description on the dataset page: https://huggingface.co/datasets/bfshi/HLVid.

sourceHugging Faceupdated 6mo agoView on Hugging Face
1likes237downloads
filevideos_part_0001.tar7.75 GBdownload
filevideos_part_0002.tar6.96 GBdownload
filevideos_part_0003.tar9.06 GBdownload
filevideos_part_0004.tar9.25 GBdownload
filevideos_part_0005.tar9.66 GBdownload
filevideos_part_0006.tar7.12 GBdownload
filevideos_part_0007.tar9.33 GBdownload
filevideos_part_0008.tar7.84 GBdownload
filevideos_part_0009.tar9.39 GBdownload
filevideos_part_0010.tar8.52 GBdownload
filevideos_part_0011.tar8.43 GBdownload
filevideos_part_0012.tar8.34 GBdownload
filevideos_part_0013.tar9.50 GBdownload
filevideos_part_0014.tar9.76 GBdownload
filevideos_part_0015.tar9.53 GBdownload
filevideos_part_0016.tar9.42 GBdownload

bfshi/HLVid · main · files are served by the source, never re-hosted here