turhancan97/SpaRRTa
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models SpaRRTa is a synthetic benchmark that probes whether Visual Foundation Models (VFMs) โ such as DINO, DINOv2/v3, MAE, CroCo, VGGT, SPA and CLIP โ encode the spatial relations between objects in a scene, rather than only their semantic identity. ๐ Paper: arXiv:2601.11729 ๐ป Code: github.com/gmum/SpaRRTa ๐งฑ Real-world (lego) split: turhancan97/SpaRRTa-Lego ๐ฌ Attention-analysis split (images +โฆ See the full description on the dataset page: https://huggingface.co/datasets/turhancan97/SpaRRTa.
This repository belongs to turhancan97 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
