owl-owl/POVBench
POVBench Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language ModelsEMNLP 2026 Findings Project page · Code Given a sentence in which an observer says where they last saw an object — from their own point of view — a model must recover that perspective and localize the target in image space. Crucially, the observer's right is not necessarily aligned with the camera's right. Three conditions progressively reduce the amount of reasoning required:… See the full description on the dataset page: https://huggingface.co/datasets/owl-owl/POVBench.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face