etri-vilab/MultihopSpatial
[ECCV 2026] MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Models Project Page | Paper | Model Overview MultihopSpatial is a benchmark designed to evaluate whether vision-language models (VLMs) demonstrate robustness in multi-hop compositional spatial reasoning. Unlike existing benchmarks that only assess single-step spatial relations, MultihopSpatial features queries with 1 to 3 reasoning hops paired with… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/MultihopSpatial.
6950
