turhancan97/SpaRRTa
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models SpaRRTa is a synthetic benchmark that probes whether Visual Foundation Models (VFMs) β such as DINO, DINOv2/v3, MAE, CroCo, VGGT, SPA and CLIP β encode the spatial relations between objects in a scene, rather than only their semantic identity. π Paper: arXiv:2601.11729 π» Code: github.com/gmum/SpaRRTa π§± Real-world (lego) split: turhancan97/SpaRRTa-Lego π¬ Attention-analysis split (images +β¦ See the full description on the dataset page: https://huggingface.co/datasets/turhancan97/SpaRRTa.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone elseβs repository from here would need an authorised integration and the account holderβs consent, so the link goes to the source instead.
Open discussions on Hugging Face