datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Visual-Searchvisual_26k_reasoningspatial-visual-reasoning-66kVisual-Jigsawgrounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.matsci-visual-reasoning-nc
MatSciChartQ-Traces (CC BY-NC 4.0)
MatSciChartQ-Traces pairs every question in the MatSciChartQ benchmark (translorentz/materials-science-visual-bench-nc, version 1.8.1-nc) with a step-by-step reasoning trace that works from the chart pixels to the graded answer. There are 4,037 traces covering all twelve chart families and all 40 question types of the parent benchmark, one trace per question row, each shipped alongside the chart image and the exact final answer. The charts… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/matsci-visual-reasoning-nc.BLIP3o-Visual-ReasoningTurkish-Visual-Reasoning-Dataset
Turkish Visual Reasoning Dataset
The Turkish Visual Reasoning Dataset is a Turkish multimodal reasoning dataset designed to evaluate and improve the abstract reasoning capabilities of Vision-Language Models (VLMs).
It was created by adapting established visual reasoning benchmarks into Turkish and combining them with original Turkish BİLSEM preparation questions. The dataset targets challenging reasoning tasks such as logical pattern discovery, spatial reasoning, analogical… See the full description on the dataset page: https://huggingface.co/datasets/Berkesule/Turkish-Visual-Reasoning-Dataset.visual-reasoning
Visual Reasoning
Dataset Summary
visual_reasoning is a visual-question-answering benchmark for probing vision-language
models (VLMs) on five core visual reasoning skills: color identification, counting,
direction (orientation), spatial relation, and shape identification. Each task has
up to three splits:
simple — procedurally generated synthetic images (basic shapes/colors), with the
object category named in the prompt.
simple_noprior — the same synthetic images… See the full description on the dataset page: https://huggingface.co/datasets/kunwang0129/visual-reasoning.Visual_ReasoningPersian-Visual-Abstraction-Reasoningvisual_text_reasoning
