datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EpiBench-NeurIPS2026
EpiBench
Anonymous release for NeurIPS 2026 Evaluations & Datasets Track review (paper ID 1899). All methodology, ablations, and analyses are in the companion paper; this card lists only what reviewers and downstream users need to load the data.
A 25,737-patient ILAE-aligned multimodal epilepsy benchmark derived from PubMed Central case reports + 192 EpiRAG textbook vignettes.
6 tasks: epilepsy_type, seizure_type, ez_localization, aed_response, surgery_outcome, status_epilepticus… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-1899-ED-2026/EpiBench-NeurIPS2026.benchmark
EditJudge-Bench
EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as
automated judges for image-edit verification. Each row contains a source image,
an edited image, a factual edit instruction, counterfactual instructions, and
ground-truth scene parameters produced by a controlled Blender/Infinigen
generation pipeline.
This repository is an anonymous review release for a NeurIPS Evaluations and
Datasets submission.
Dataset Contents
1… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.birdsnap_liteThis is a version of BirdSnap that will be easier on your free Google Colab quota.
