bfshi/HLVid
HLVid Dataset Project Page | Paper | GitHub HLVid (High-resolution, Long-form Video QA) is a benchmark introduced in the paper "Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing". It is designed to evaluate Multi-modal Large Language Models (MLLMs) on long-form, high-resolution video understanding. The benchmark features 5-minute videos at 4K resolution, challenging models to handle significant spatiotemporal redundancy while… See the full description on the dataset page: https://huggingface.co/datasets/bfshi/HLVid.
1237
videos_part_0001.tardownload
videos_part_0002.tardownload
videos_part_0003.tardownload
videos_part_0004.tardownload
videos_part_0005.tardownload
videos_part_0006.tardownload
videos_part_0007.tardownload
videos_part_0008.tardownload
videos_part_0009.tardownload
videos_part_0010.tardownload
videos_part_0011.tardownload
videos_part_0012.tardownload
videos_part_0013.tardownload
videos_part_0014.tardownload
videos_part_0015.tardownload
videos_part_0016.tardownload
