CoolFace
Datasetpublic

turhancan97/SpaRRTa

SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models SpaRRTa is a synthetic benchmark that probes whether Visual Foundation Models (VFMs) โ€” such as DINO, DINOv2/v3, MAE, CroCo, VGGT, SPA and CLIP โ€” encode the spatial relations between objects in a scene, rather than only their semantic identity. ๐Ÿ“„ Paper: arXiv:2601.11729 ๐Ÿ’ป Code: github.com/gmum/SpaRRTa ๐Ÿงฑ Real-world (lego) split: turhancan97/SpaRRTa-Lego ๐Ÿ”ฌ Attention-analysis split (images +โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/turhancan97/SpaRRTa.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
1likes353downloads

turhancan97/SpaRRTa ยท main ยท files are served by the source, never re-hosted here