neurips-2026
EpiBench-NeurIPS2026
EpiBench
Anonymous release for NeurIPS 2026 Evaluations & Datasets Track review (paper ID 1899). All methodology, ablations, and analyses are in the companion paper; this card lists only what reviewers and downstream users need to load the data.
A 25,737-patient ILAE-aligned multimodal epilepsy benchmark derived from PubMed Central case reports + 192 EpiRAG textbook vignettes.
6 tasks: epilepsy_type, seizure_type, ez_localization, aed_response, surgery_outcome, status_epilepticus… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-1899-ED-2026/EpiBench-NeurIPS2026.NeurIPS_2026_BLV
BLV Object Recognition: Synthetic + Real-World
A dataset for training and evaluating object recognition and segmentation
models on infrastructure relevant to blind and low-vision (BLV) navigation
in urban environments. Three configurations plus a flat tree of 3D assets:
Config / tree
Splits
Purpose
syn
train
Photorealistic IsaacSim renders for training / pretraining.
real_ours
train / validation / test
Real photographs we captured. real_ours/test is the canonical… See the full description on the dataset page: https://huggingface.co/datasets/NavAble/NeurIPS_2026_BLV.neurips-2026-evals
NeurIPS 2026 Agent Evaluation Dataset
This dataset contains evaluation results for various AI agents across multiple benchmarks.
Dataset Structure
The dataset is organized by model (as configs) with each benchmark as a split.
Each model/benchmark folder contains:
Main results file (.jsonl or .parquet format)
Summary statistics (.summary.json) - for models with metadata
Configuration file (.toml) - for models with metadata
Traces folder with execution traces… See the full description on the dataset page: https://huggingface.co/datasets/akenginorhun/neurips-2026-evals.IGF-Bench
IGF-Bench: Indoor Geometric Fidelity Benchmark
Anonymous mirror for NeurIPS 2026 Evaluations and Datasets Track double-blind review.
The de-anonymised author/maintainer information will replace this header at camera-ready.
IGF-Bench is the first benchmark for evaluating structural-level geometric fidelity of conditionally generated indoor scene images, going beyond perceptual metrics like FID and LPIPS. It pairs 3,600 calibrated synthetic ground-truth views with 21,600 generated… See the full description on the dataset page: https://huggingface.co/datasets/igfbench-neurips2026/IGF-Bench.PRISM-Dataset
PRISM: Polarimetric Road-surface Intelligent Sensing and Measurement Dataset
Anonymous submission to NeurIPS 2026 Evaluations & Datasets Track.
Author identities and the camera-ready release URL will be revealed at the
camera-ready stage.
PRISM is a polarimetric road-surface dataset and benchmark of 47,098
time-synchronized frames combining trichromatic linear polarization,
co-boresighted RGB, 128-channel LiDAR, and RTK-GNSS/INS, captured from an
in-vehicle sensor rig across… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-2026-PRISM/PRISM-Dataset.ARES-Bench
ARES-Bench
ARES-Bench is the open audit substrate released with the paper
Auditing LLM User Simulators for Recommender A/B Testing (NeurIPS 2026, ED
Track, under review). It turns the ARES reliability-audit view — the LLM
backbone is the measurement instrument under test, not an interchangeable
implementation detail — into a reproducible protocol over structured
behavioral logs, a portable visual sandbox, and a screenshot cache.
This release hosts the 17,000-session core corpus that… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-ares-authors/ARES-Bench.
