Alibaba-NLP/xvbench
XVBench XVBench is a benchmark for evaluating multimodal retrieval-augmented generation systems on cross-video understanding. It is introduced alongside the paper VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph. The questions in XVBench are created based on videos from HowTo100M, a large-scale corpus of narrated instructional videos. The benchmark focuses on questions that require models or agents to retrieve and reason… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/xvbench.
037
update
update
dataset initial
initial commit
