CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Anonymous1383 /ship-dataset ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning ShipBench is a metadata-grounded vision-language benchmark on parametrically-generated ship structural drawings. Six commercial ship types × nine drawing-grounded sub-tasks × deterministic ground truth derived directly from the generator's input dictionary (no human annotation, no rule-citation labels). Quick reference Total candidates: 6{,}450 across 6 ship types (Tanker, VLCC, BULKC, CNTR, LNGC… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1383/ship-dataset.imagevisual-question-answering10K<n<100K0 likes638 downloads5mo agoHugging Face02anonymous-8421 /VL-DocIRgated Abstract VL-DocIR is a page-level benchmark for vision-based long-document retrieval built from 29,641 documents rendered into 388,548 page images from Wikipedia, arXiv, PubMed, and SEC proxy statements. The benchmark contains 271,760 questions over 23 domains and six query types, covering single-page, multi-page, and cross-document evidence configurations. Questions are grounded to rendered pages and HTML element identifiers, then filtered with a cleaning pipeline that targets… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-8421/VL-DocIR.textvisual-document-retrieval1M<n<10M0 likes281 downloads5mo agoHugging Face03AnonymousNu /SciRec SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction This dataset contains multimodal question-answering examples grounded in textbook figures. Records in the figure-grounded configurations are filtered to include only examples whose referenced image files are present in this release. Configurations visual: 13791 figure-grounded visual questions with resolved images. knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousNu/SciRec.imagevisual-question-answering10K<n<100K0 likes135 downloads5mo agoHugging Face04Anonymousblind /agent-failure-dynamics AgentHazard: Process-Centric Benchmark for AI Agent Trajectory Analysis AgentHazard is the first benchmark designed specifically for process-level analysis of AI coding agent trajectories. Unlike existing benchmarks that evaluate only final outcomes (pass/fail), AgentHazard provides standardized edit-level annotations, hazard estimation protocols, and stopping-policy evaluation tasks with unified evaluation across 85,050 trajectories from 6+ agent families. Why… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousblind/agent-failure-dynamics.documenttext-classification1K<n<10K0 likes126 downloads3mo agoHugging Face05anonymous80934 /HalluCompass HalluCompass A direction-aware diagnostic benchmark and protocol for vision-language model (VLM) hallucination evaluation NeurIPS 2026 Datasets & Benchmarks Track Unified release: 2,200 images across MS-COCO + AMBER + NoCaps + VizWiz, 10,000 queries, two-pass annotation (GPT-4o-mini 1st-pass + 5-human verification, Fleiss' κ = 0.72 inter-annotator agreement on a stratified 250-image validation subset, substantial agreement per Landis & Koch 1977). A balanced POPE-compatible… See the full description on the dataset page: https://huggingface.co/datasets/anonymous80934/HalluCompass.imagevisual-question-answering10K<n<100K0 likes92 downloads5mo agoHugging Face06anonymous-submission-RIG-bench /RIG-Bench RIG-bench Anonymous submission to the NeurIPS 2026 Evaluations & Datasets (E&D) Track. A benchmark for reasoning-driven image generation: given visual context (images + instruction + optional demonstration pairs), the model must produce the answer as a single image. 2,000 samples 4 task families × 11 subtasks ~1.4 GB Files RIG-bench/ ├── README.md ├── samples.jsonl # 2,000 records └── images/<sample_id>/ ├── input_<order>.<ext> ├──… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-RIG-bench/RIG-Bench.imageimage-to-image1K<n<10K0 likes88 downloads5mo agoHugging Face07kir0504 /anonymous-clear-drive CLEAR-Drive Review Sample This repository provides an anonymized review-only sample of CLEAR-Drive, an autonomous-driving causal reasoning dataset introduced in the submitted paper. CLEAR-Drive is designed to improve the logical judgment ability of driving-oriented vision-language models. It constructs positive--negative sample pairs from Chain-of-Causality (CoC) reasoning, where the positive sample represents a causally valid reasoning trace and the negative sample is produced by… See the full description on the dataset page: https://huggingface.co/datasets/kir0504/anonymous-clear-drive.image1K<n<10K0 likes39 downloads5mo agoHugging Face08WorldModelBenchmark /playworld-bench-anonymous PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Anonymous benchmark release for double-blind review. Author, affiliation, and identity-bearing links are intentionally omitted. Project | Code Video world models generate future states from an initial observation and user actions. Comparing interactive models fairly is difficult because the actions required to reach the same long-horizon objective can vary substantially across models. PlayWorld… See the full description on the dataset page: https://huggingface.co/datasets/WorldModelBenchmark/playworld-bench-anonymous.imagevisual-question-answeringn<1K0 likes38 downloads9h agoHugging Face09QuantumWhisper42 /anonymous_dataset KnowVis: A Dual-View Benchmark for Diagnosing World-Knowledge Grounding in Text-to-Image Models 📖 Overview Text-to-image (T2I) models have made substantial progress in visual realism, aesthetic quality, and instruction following. However, real-world prompts often go beyond explicit visual descriptions and require implicit facts, structured knowledge, and domain-specific commonsense. Existing evaluations mainly focus on explicit prompt-to-image semantic alignment or… See the full description on the dataset page: https://huggingface.co/datasets/QuantumWhisper42/anonymous_dataset.imagetext-to-image1K<n<10K0 likes31 downloads5mo agoHugging Face10anonymous-mY2nG5 /H3DBenchHierarchical 3D Benmark This is an annotation dataset for 3D quality evaluation, including Object-Level, Part-Level and Material-Subject annotations. We also release 3D assets generated from new 3D generative models that are not included in 3DGen-Bench dataset. 3dtext-to-3dn<1K0 likes29 downloads1y agoHugging Face11Anonymouszzzz /PuzzleCodeBench Dataset Card Overview This dataset contains the anchor split for public browsing and experimentation. Transfer Split transfer.jsonl is the hidden evaluation split. It is used to run generated solver code on unseen instances for benchmark evaluation, and is intentionally not exposed in the public viewer split configuration. license: cc-by-4.0 imagen<1K0 likes19 downloads5mo agoHugging Face12anonymous38304 /DriveSpatialimage10K<n<100K0 likes13 downloads5mo agoHugging Face13Anonymous02345 /ChallengeBench ChallengeBench ChallengeBench is a high-difficulty diagnostic benchmark for evaluating and analyzing multimodal large language models (MLLMs). It is introduced as the diagnostic substrate of ErrorInsight, a multimodal agent framework for structured, case-level error diagnosis of MLLMs. Unlike conventional benchmarks that mainly report aggregate accuracy, ChallengeBench is designed to expose model failure boundaries and provide structured task labels and probe-question information… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous02345/ChallengeBench.imagen<1K0 likes7 downloads5mo agoHugging Face14anonymous-8421 /VL-DocIR-RepresentativeSubsetgated Representative Subset Creation This dataset represents a representative subset of VL-DocIR dataset (https://huggingface.co/datasets/anonymous-8421/VL-DocIR). Subset creation method: Random sampling of 5 queries for the combination of each data source and evidence structure type (we ensure that no query is taken more than once). This results in 120 queries. Collection of all documents that are referenced by query evidence. Abstract VL-DocIR is a page-level benchmark… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-8421/VL-DocIR-RepresentativeSubset.textvisual-document-retrieval1K<n<10K0 likes6 downloads5mo agoHugging Face15kapididi /ICLR26_anonymousimage1K<n<10K0 likes4 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.