datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpineBench
Dataset Card for SpineBench
Benchmark Details
Paper Information
Benchmark Examples
Benchmark Distribution
Data Format
Data SourceHuman Evaluation of MLLMs Reasoning Performance
Citation
Benchmark Details
SpineBench is a comprehensive Visual Question Answering (VQA) benchmark designed for fine-grained analysis and evaluation of LVLM in the spinal domain. SpineBench comprises 64,878 QA pairs from 40,263 spine images, covering 11 spinal diseases through two critical… See the full description on the dataset page: https://huggingface.co/datasets/Silversorrow/SpineBench.SpinBench
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
🌐 Project page |
🤗 Dataset |
📑 Paper |
💻 Code
SpinBench is a cognitively grounded diagnostic benchmark for evaluating spatial reasoning in vision-language models (VLMs).
SpinBench is designed around the core challenge of spatial reasoning: perspective taking, the ability to reason about how scenes and object relations change under viewpoint transformation. Since perspective taking requires… See the full description on the dataset page: https://huggingface.co/datasets/YuyouZhang/SpinBench.CelebaControlnet
