datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Anonymous_ACMMM_2025_Submission
🗂️ Anonymous_ACMMM_2025_Submission Dataset
This dataset is prepared for the Anonymous ACMMM 2025 submission, containing multi-view event-based data designed for dynamic 3D scene reconstruction tasks.
📁 Dataset Structure
Each subfolder corresponds to a distinct synthetic or real-world scene, such as:
lego_6_views/
capsule_6_views/
garage_6_views/
Restroom_6_views/
Cubes_6_views/
Hinge_6_views/
MC-Toy_6_views/
Rubik’s-Cube_6_views/
Each scene folder contains 6 views… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-ACMMM-2025-Submission/Anonymous_ACMMM_2025_Submission.CaReCoS
CaReCoS
A medical acoustic question-answering dataset for reasoning over mel spectrograms
of heart, lung, and cough sounds. Each record provides a clinical question, the
mel-spectrogram image of a recording, a ground-truth answer, and the
recording's clinical metadata.
The task is purely visual: a model receives the spectrogram image together with the
question and must reason over the spectrogram to produce the answer. The raw audio is
not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.text2tile_large
Dataset Card for "text2tile_large"
More Information needed
TiBuDBRIG-Bench-Qualitative
Extended Qualitative Examples for RIG-Bench
This page provides additional qualitative examples complementing the paper. It is created in response to reviewer feedback requesting more generated examples illustrating both successful reasoning and common failure modes across model families.
These examples visualize the RIG-Bench setting: a model receives visual context and a short instruction, infers the latent rule or target outcome, and synthesizes the answer directly as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-RIG-bench/RIG-Bench-Qualitative.RIG-Bench
RIG-bench
Anonymous submission to the NeurIPS 2026 Evaluations & Datasets (E&D) Track.
A benchmark for reasoning-driven image generation: given visual context (images + instruction + optional demonstration pairs), the model must produce the answer as a single image.
2,000 samples
4 task families × 11 subtasks
~1.4 GB
Files
RIG-bench/
├── README.md
├── samples.jsonl # 2,000 records
└── images/<sample_id>/
├── input_<order>.<ext>
├──… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-RIG-bench/RIG-Bench.Poison-3DGS
Poison-3DGS (review copy)
Anonymous data release accompanying a paper under double-blind review.
For review purposes only. Please do not redistribute it or use it for any other purpose.
This review set contains 150 scenes (125 poisoned and 25 clean), built from 25 base scenes. The full
benchmark of 1,353 scenes (1,253 poisoned and 100 clean) will be released after acceptance. The release
contains poisoned training data and manipulated 3D Gaussian Splatting models.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-user-submission/Poison-3DGS.bird3m
Bird3M Dataset
Dataset Description
Bird3M is the first synchronized, multi-modal, multi-individual dataset designed for comprehensive behavioral analysis of freely interacting birds, specifically zebra finches, in naturalistic settings. It addresses the critical need for benchmark datasets that integrate precisely synchronized multi-modal recordings to support tasks such as 3D pose estimation, multi-animal tracking, sound source localization, and vocalization attribution.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission000/bird3m.InspectVQA
InspectVQA
InspectVQA is an expert-grounded visual question answering dataset for subsea pipe inspection. It contains underwater industrial inspection images, multi-label surface condition annotations, optional segmentation masks, and inspection-oriented question-answer pairs.
Dataset structure
images/: inspection images.
gt/: segmentation masks where available.
annotations/label.jsonl: original annotation file.
dataset_statistics.json: dataset statistics.… See the full description on the dataset page: https://huggingface.co/datasets/anonymousSubmissionVqa2026/InspectVQA.OTA-76k
Dataset Card for OTA-76k (POIROT Framework)
Dataset Summary
OTA-76k is a large-scale, bounding-box-grounded, multi-step video reasoning dataset designed to train Multimodal Large Language Models (MLLMs) for fine-grained, spatio-temporal deduction. This dataset addresses the common pitfalls of existing models in video reasoning, such as their over-reliance on frame-level perception and outcome-oriented sparse rewards, which often lead to visual noise interference… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-221/OTA-76k.tactile-mnist-touch-real-single-t256-320x240Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
TextureBench-3DGS
TextureBench-3DGS
Note. This is an anonymised copy of the benchmark released for the double-blind review process. It contains no author
information; the full release with attribution, licence details and accompanying code will be published after the review.
TextureBench-3DGS is a benchmark of 28 real outdoor scenes with dense natural texture (gravel, grass, leaves, brick, concrete,
tiles, foliage), each captured as a multi-view photo set suitable for 3D Gaussian Splatting… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-user-submission/TextureBench-3DGS.RFSchemBench
RFSchemBench
A multimodal LLM evaluation benchmark for radio-frequency circuit schematic understanding, organized by a four-level semantic hierarchy:
Component Understanding — visible component, parameter, label, and supply-rail recognition.
Structural Understanding — net membership, pin-to-net mapping, boundary connectivity, and pair-via-net topological reasoning.
Functional Understanding — circuit functional role, signal-form classification, supply strategy, sub-type… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission042/RFSchemBench.tactile-mnist-touch-syn-single-t32-64x64Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
tactile-mnist-touch-syn-single-t32-320x240Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
pixels_vs_code
Pattern Over Pixels Screenshot-to-Code
This dataset contains controlled counterfactual screenshot-to-code examples built
from 30 real-world webpages from Design2Code. Each example preserves a repeated
UI pattern while introducing a single localized deviation, allowing researchers
to test whether multimodal models follow the pixels or simply restore the
dominant template.
Contents
720 perturbed HTML instances
360 structural-card examples
360 text-style examples
2… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSubmissionASE/pixels_vs_code.tactile-mnist-touch-starstruck-syn-single-t32-64x64Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
mmscibench-anonymous
Dataset Name
Dataset Summary
MMSciBench dataset focuses on mathematics and physics that evaluates scientific reasoning capabilities.
Dataset Structure
Data Instances
QA_has_img_with_categories.csv: Contains data of Q&A questions (with images) of the subject indicated by the folder name.
QA_no_img_with_categories.csv: Contains data of text-only Q&A questions of the subject indicated by the folder name.
multi_choice_has_img_with_categories.csv:… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl-submission/mmscibench-anonymous.tactile-mnist-touch-real-single-t256-64x64Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
tactile-mnist-touch-starstruck-syn-single-t32-320x240Documentation is available at https://github.com/[REDACTED]/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
CausalConflictBench
