datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ship-dataset
ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning
ShipBench is a metadata-grounded vision-language benchmark on parametrically-generated ship structural drawings. Six commercial ship types × nine drawing-grounded sub-tasks × deterministic ground truth derived directly from the generator's input dictionary (no human annotation, no rule-citation labels).
Quick reference
Total candidates: 6{,}450 across 6 ship types (Tanker, VLCC, BULKC, CNTR, LNGC… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1383/ship-dataset.VL-DocIR
Abstract
VL-DocIR is a page-level benchmark for vision-based long-document retrieval built from 29,641 documents rendered into 388,548 page images from Wikipedia, arXiv, PubMed, and SEC proxy statements. The benchmark contains 271,760 questions over 23 domains and six query types, covering single-page, multi-page, and cross-document evidence configurations. Questions are grounded to rendered pages and HTML element identifiers, then filtered with a cleaning pipeline that targets… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-8421/VL-DocIR.SciRec
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousNu/SciRec.agent-failure-dynamics
AgentHazard: Process-Centric Benchmark for AI Agent Trajectory Analysis
AgentHazard is the first benchmark designed specifically for process-level analysis of AI coding agent trajectories. Unlike existing benchmarks that evaluate only final outcomes (pass/fail), AgentHazard provides standardized edit-level annotations, hazard estimation protocols, and stopping-policy evaluation tasks with unified evaluation across 85,050 trajectories from 6+ agent families.
Why… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousblind/agent-failure-dynamics.HalluCompass
HalluCompass
A direction-aware diagnostic benchmark and protocol for vision-language model (VLM) hallucination evaluation
NeurIPS 2026 Datasets & Benchmarks Track
Unified release: 2,200 images across MS-COCO + AMBER + NoCaps + VizWiz, 10,000 queries, two-pass annotation (GPT-4o-mini 1st-pass + 5-human verification, Fleiss' κ = 0.72 inter-annotator agreement on a stratified 250-image validation subset, substantial agreement per Landis & Koch 1977). A balanced POPE-compatible… See the full description on the dataset page: https://huggingface.co/datasets/anonymous80934/HalluCompass.RIG-Bench
RIG-bench
Anonymous submission to the NeurIPS 2026 Evaluations & Datasets (E&D) Track.
A benchmark for reasoning-driven image generation: given visual context (images + instruction + optional demonstration pairs), the model must produce the answer as a single image.
2,000 samples
4 task families × 11 subtasks
~1.4 GB
Files
RIG-bench/
├── README.md
├── samples.jsonl # 2,000 records
└── images/<sample_id>/
├── input_<order>.<ext>
├──… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-RIG-bench/RIG-Bench.anonymous-clear-drive
CLEAR-Drive Review Sample
This repository provides an anonymized review-only sample of CLEAR-Drive, an autonomous-driving causal reasoning dataset introduced in the submitted paper.
CLEAR-Drive is designed to improve the logical judgment ability of driving-oriented vision-language models. It constructs positive--negative sample pairs from Chain-of-Causality (CoC) reasoning, where the positive sample represents a causally valid reasoning trace and the negative sample is produced by… See the full description on the dataset page: https://huggingface.co/datasets/kir0504/anonymous-clear-drive.playworld-bench-anonymous
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
Anonymous benchmark release for double-blind review. Author, affiliation, and
identity-bearing links are intentionally omitted.
Project | Code
Video world models generate future states from an initial observation and user
actions. Comparing interactive models fairly is difficult because the actions
required to reach the same long-horizon objective can vary substantially across
models.
PlayWorld… See the full description on the dataset page: https://huggingface.co/datasets/WorldModelBenchmark/playworld-bench-anonymous.anonymous_dataset
KnowVis: A Dual-View Benchmark for Diagnosing World-Knowledge Grounding in Text-to-Image Models
📖 Overview
Text-to-image (T2I) models have made substantial progress in visual realism, aesthetic quality, and instruction following. However, real-world prompts often go beyond explicit visual descriptions and require implicit facts, structured knowledge, and domain-specific commonsense. Existing evaluations mainly focus on explicit prompt-to-image semantic alignment or… See the full description on the dataset page: https://huggingface.co/datasets/QuantumWhisper42/anonymous_dataset.H3DBenchHierarchical 3D Benmark
This is an annotation dataset for 3D quality evaluation, including Object-Level, Part-Level and Material-Subject annotations.
We also release 3D assets generated from new 3D generative models that are not included in 3DGen-Bench dataset.
PuzzleCodeBench
Dataset Card
Overview
This dataset contains the anchor split for public browsing and experimentation.
Transfer Split
transfer.jsonl is the hidden evaluation split. It is used to run generated solver code on unseen instances for benchmark evaluation, and is intentionally not exposed in the public viewer split configuration.
license: cc-by-4.0
DriveSpatialChallengeBench
ChallengeBench
ChallengeBench is a high-difficulty diagnostic benchmark for evaluating and analyzing multimodal large language models (MLLMs). It is introduced as the diagnostic substrate of ErrorInsight, a multimodal agent framework for structured, case-level error diagnosis of MLLMs.
Unlike conventional benchmarks that mainly report aggregate accuracy, ChallengeBench is designed to expose model failure boundaries and provide structured task labels and probe-question information… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous02345/ChallengeBench.VL-DocIR-RepresentativeSubset
Representative Subset Creation
This dataset represents a representative subset of VL-DocIR dataset (https://huggingface.co/datasets/anonymous-8421/VL-DocIR).
Subset creation method:
Random sampling of 5 queries for the combination of each data source and evidence structure type (we ensure that no query is taken more than once).
This results in 120 queries.
Collection of all documents that are referenced by query evidence.
Abstract
VL-DocIR is a page-level benchmark… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-8421/VL-DocIR-RepresentativeSubset.ICLR26_anonymous
