datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EpiBench-NeurIPS2026
EpiBench
Anonymous release for NeurIPS 2026 Evaluations & Datasets Track review (paper ID 1899). All methodology, ablations, and analyses are in the companion paper; this card lists only what reviewers and downstream users need to load the data.
A 25,737-patient ILAE-aligned multimodal epilepsy benchmark derived from PubMed Central case reports + 192 EpiRAG textbook vignettes.
6 tasks: epilepsy_type, seizure_type, ez_localization, aed_response, surgery_outcome, status_epilepticus… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-1899-ED-2026/EpiBench-NeurIPS2026.NeurIPS_2026_BLV
BLV Object Recognition: Synthetic + Real-World
A dataset for training and evaluating object recognition and segmentation
models on infrastructure relevant to blind and low-vision (BLV) navigation
in urban environments. Three configurations plus a flat tree of 3D assets:
Config / tree
Splits
Purpose
syn
train
Photorealistic IsaacSim renders for training / pretraining.
real_ours
train / validation / test
Real photographs we captured. real_ours/test is the canonical… See the full description on the dataset page: https://huggingface.co/datasets/NavAble/NeurIPS_2026_BLV.IGF-Bench
IGF-Bench: Indoor Geometric Fidelity Benchmark
Anonymous mirror for NeurIPS 2026 Evaluations and Datasets Track double-blind review.
The de-anonymised author/maintainer information will replace this header at camera-ready.
IGF-Bench is the first benchmark for evaluating structural-level geometric fidelity of conditionally generated indoor scene images, going beyond perceptual metrics like FID and LPIPS. It pairs 3,600 calibrated synthetic ground-truth views with 21,600 generated… See the full description on the dataset page: https://huggingface.co/datasets/igfbench-neurips2026/IGF-Bench.ARES-Bench
ARES-Bench
ARES-Bench is the open audit substrate released with the paper
Auditing LLM User Simulators for Recommender A/B Testing (NeurIPS 2026, ED
Track, under review). It turns the ARES reliability-audit view — the LLM
backbone is the measurement instrument under test, not an interchangeable
implementation detail — into a reproducible protocol over structured
behavioral logs, a portable visual sandbox, and a screenshot cache.
This release hosts the 17,000-session core corpus that… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-ares-authors/ARES-Bench.PRISM-Dataset
PRISM: Polarimetric Road-surface Intelligent Sensing and Measurement Dataset
Anonymous submission to NeurIPS 2026 Evaluations & Datasets Track.
Author identities and the camera-ready release URL will be revealed at the
camera-ready stage.
PRISM is a polarimetric road-surface dataset and benchmark of 47,098
time-synchronized frames combining trichromatic linear polarization,
co-boresighted RGB, 128-channel LiDAR, and RTK-GNSS/INS, captured from an
in-vehicle sensor rig across… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-2026-PRISM/PRISM-Dataset.NeurIPS2026-DermVerse-500K
DermVerse-500K
Paper: DermVerse-500K: A Large-Scale Expert-Annotated Dataset with VLM Assisted Clinical Feature ExtractionSubmitted to: 40th Conference on Neural Information Processing Systems (NeurIPS 2026)Paper ID: 4127
Dataset Summary
DermVerse-500K is a large-scale dermatology dataset comprising 500,000 clinical images paired with fine-grained, structured clinical descriptors. It is designed to address critical limitations in existing dermatology AI resources:… See the full description on the dataset page: https://huggingface.co/datasets/BBAnonymous/NeurIPS2026-DermVerse-500K.tdbench-review
TDBench: Benchmarking Vision-Language Models on Top-Down Images
Note (Anonymous Review Version). This dataset card accompanies a NeurIPS 2026 Evaluations & Datasets double-blind submission. Author identifiers, institutional affiliations, project pages, and external repository links have been removed for the review period. The full set of public artifacts and the final citation will be restored upon decision.
Overview
TDBench is a benchmark for evaluating… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-tdbench/tdbench-review.NeurIPS_2026_BLV_Subset
4 GB stratified preview. Full dataset: NavAble/NeurIPS_2026_BLV.
BLV Object Recognition: Synthetic + Real-World
A dataset for training and evaluating object recognition and segmentation
models on infrastructure relevant to blind and low-vision (BLV) navigation
in urban environments. Three configurations plus a flat tree of 3D assets:
Config / tree
Splits
Purpose
syn
train
Photorealistic IsaacSim renders for training / pretraining.
real_ours
train / validation / test… See the full description on the dataset page: https://huggingface.co/datasets/NavAble/NeurIPS_2026_BLV_Subset.PRISM-Dataset-Sample
PRISM Sample: Polarimetric Road-surface Intelligent Sensing and Measurement Dataset
Anonymous submission to NeurIPS 2026 Evaluations & Datasets Track.
This is a representative sample of the PRISM dataset, designed to enable reviewers and researchers to inspect data quality without downloading the full ~1.6 TB dataset.
Why a sample dataset?
The full PRISM dataset contains 47,098 time-synchronized frames across 41 sessions. This sample provides:
Quick quality inspection:… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-2026-PRISM/PRISM-Dataset-Sample.signalbench-openapps
SignalBench OpenApps Dense Signal Dataset
OpenApps browser-task evaluation points, Set-of-Marks screenshots, and Monte Carlo dense-signal labels under a scripted policy.
This repository contains one SignalBench dataset with two synchronized views:
runtime/dataset.pkl is the executable artifact used by the SignalBench benchmark code.
data/examples.parquet has exactly one row per benchmark example, with
state/action/next-state text, renderable state_image and
next_state_image columns… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-anonymous/signalbench-openapps.signalbench-alfworld-bkp
SignalBench ALFWorld Dense Signal Dataset
ALFWorld household-task evaluation points, visual state observations, and Monte Carlo dense-signal labels under a scripted policy.
This repository contains one SignalBench dataset with two synchronized views:
runtime/dataset.pkl is the executable artifact used by the SignalBench benchmark code.
data/examples.parquet has exactly one row per benchmark example, with
state/action/next-state text, renderable state_image and
next_state_image… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-anonymous/signalbench-alfworld-bkp.signalbench-frozenlake-bkp
SignalBench FrozenLake Dense Signal Dataset
FrozenLake 8x8 evaluation points, rendered state images, and Monte Carlo dense-signal labels under a scripted policy.
This repository contains one SignalBench dataset with two synchronized views:
runtime/dataset.pkl is the executable artifact used by the SignalBench benchmark code.
data/examples.parquet has exactly one row per benchmark example, with
state/action/next-state text, renderable state_image and
next_state_image columns when… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-anonymous/signalbench-frozenlake-bkp.signalbench-openapps-bkp
SignalBench OpenApps Dense Signal Dataset
OpenApps browser-task evaluation points, Set-of-Marks screenshots, and Monte Carlo dense-signal labels under a scripted policy.
This repository contains one SignalBench dataset with two synchronized views:
runtime/dataset.pkl is the executable artifact used by the SignalBench benchmark code.
data/examples.parquet has exactly one row per benchmark example, with
state/action/next-state text, renderable state_image and
next_state_image columns… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-anonymous/signalbench-openapps-bkp.signalbench-alfworld
SignalBench ALFWorld Dense Signal Dataset
ALFWorld household-task evaluation points, visual state observations, and Monte Carlo dense-signal labels under a scripted policy.
This repository contains one SignalBench dataset with two synchronized views:
runtime/dataset.pkl is the executable artifact used by the SignalBench benchmark code.
data/examples.parquet has exactly one row per benchmark example, with
state/action/next-state text, renderable state_image and
next_state_image… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-anonymous/signalbench-alfworld.signalbench-frozenlake
SignalBench FrozenLake Dense Signal Dataset
FrozenLake 8x8 evaluation points, rendered state images, and Monte Carlo dense-signal labels under a scripted policy.
This repository contains one SignalBench dataset with two synchronized views:
runtime/dataset.pkl is the executable artifact used by the SignalBench benchmark code.
data/examples.parquet has exactly one row per benchmark example, with
state/action/next-state text, renderable state_image and
next_state_image columns when… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-anonymous/signalbench-frozenlake.PolyTopoBench
PolyTopoBench
PolyTopoBench is a remote-sensing benchmark for evaluating vector polygon generation with both exterior boundaries and interior rings. The dataset release contains two aligned sources used in the benchmark:
inria_dataset_aligned: aligned aerial imagery, building masks, and polygon annotations derived from the Inria building dataset.
deventer_512_valtest_as_val: 512 x 512 aerial image tiles with polygon annotations for multiple land-cover classes; validation and test… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS2026EDTrack/PolyTopoBench.
