datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
visual-reasoning-benchmark-results
Visual Reasoning Benchmark Suite v3.3 · 2005 Tasks · 12 Tracks Equal Weight
本版本以用户最新上传的 visual_reasoning_benchmark_suite_v3_修改 为唯一基础版本,不回退、不覆盖用户已经重绘或修改过的既有数据。完整性比对结果:原基础包中 3283 个既有数据文件全部保持字节级不变。
在此基础上新增并整合:
Nonogram(数织)150 题:45 Easy / 60 Medium / 45 Hard;
Tangram(七巧板)150 题:45 Easy / 60 Medium / 45 Hard;
两个任务的一键生成器、统一生成入口、统一评估入口、雷达图和排行榜支持。
最终总规模:2005 题,12 个 Track。
任务与数量
Task
Count
figure_completion
394
spatial_generation
56
maze_beginner
64… See the full description on the dataset page: https://huggingface.co/datasets/songyiren/visual-reasoning-benchmark-results.VisualReasoningTracerVisual-Searchvisual_26k_reasoningVisual-Jigsawspatial-visual-reasoning-66kmatsci-visual-reasoning-nc
MatSciChartQ-Traces (CC BY-NC 4.0)
MatSciChartQ-Traces pairs every question in the MatSciChartQ benchmark (translorentz/materials-science-visual-bench-nc, version 1.8.1-nc) with a step-by-step reasoning trace that works from the chart pixels to the graded answer. There are 4,037 traces covering all twelve chart families and all 40 question types of the parent benchmark, one trace per question row, each shipped alongside the chart image and the exact final answer. The charts… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/matsci-visual-reasoning-nc.grounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.visual-spatial-reasoningThe Visual Spatial Reasoning (VSR) corpus is a collection of caption-image pairs with true/false labels. Each caption describes the spatial relation of two individual objects in the image, and a vision-language model (VLM) needs to judge whether the caption is correctly describing the image (True) or not (False).BLIP3o-Visual-Reasoningvisual-spatial-reasoningThe Visual Spatial Reasoning (VSR) corpus is a collection of caption-image pairs with true/false labels. Each caption describes the spatial relation of two individual objects in the image, and a vision-language model (VLM) needs to judge whether the caption is correctly describing the image (True) or not (False).VLM-CapCurriculum-VisualReasoning-Data
VLM-CapCurriculum-VisualReasoning (D_vis)
Stage-3 visual-reasoning data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
A 16,195-sample mix of visual math and figure-grounded reasoning, drawn from four open-source corpora and packed alongside the source images. Every row also ships with a precomputed pass_rate so the same data can be ordered by sample difficulty for… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-VisualReasoning-Data.Turkish-Visual-Reasoning-Dataset
Turkish Visual Reasoning Dataset
The Turkish Visual Reasoning Dataset is a Turkish multimodal reasoning dataset designed to evaluate and improve the abstract reasoning capabilities of Vision-Language Models (VLMs).
It was created by adapting established visual reasoning benchmarks into Turkish and combining them with original Turkish BİLSEM preparation questions. The dataset targets challenging reasoning tasks such as logical pattern discovery, spatial reasoning, analogical… See the full description on the dataset page: https://huggingface.co/datasets/Berkesule/Turkish-Visual-Reasoning-Dataset.visual-reasoning
Visual Reasoning
Dataset Summary
visual_reasoning is a visual-question-answering benchmark for probing vision-language
models (VLMs) on five core visual reasoning skills: color identification, counting,
direction (orientation), spatial relation, and shape identification. Each task has
up to three splits:
simple — procedurally generated synthetic images (basic shapes/colors), with the
object category named in the prompt.
simple_noprior — the same synthetic images… See the full description on the dataset page: https://huggingface.co/datasets/kunwang0129/visual-reasoning.Visual_Reasoningvisual-logic-reasoning-v1Persian-Visual-Abstraction-Reasoningvisual-spatial-reasoning-sampled-v2visual_text_reasoningMultimodal-Visual-Reasoning-Datasetvisual-reasoning-dataset-full
