datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fragbench
FragBench (Public Tier)
Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track.
Author identity will be revealed at camera-ready.
Dataset Summary
FragBench is a benchmark for evaluating cross-session, fragmented attacks on
LLM agents that use tools via the Model Context Protocol (MCP). Each campaign
is decomposed into many small fragments distributed across sessions; a
defender must reconstruct the compositional intent. The public tier in this
repository… See the full description on the dataset page: https://huggingface.co/datasets/anon-fragbench-neurips/fragbench.adaptive-adversaries-data
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
A 21-scenario multi-turn (15-round) adversarial red-teaming benchmark for LLM agents, in which both attacker and defender are independent LLM agents and attacks are regenerated per battle. Includes calibrated 3×3 attacker × defender matrix evaluation, full battle transcripts, attack-replay corpus, and traces from two open AgentBeats competitions.
Companion paper: Adaptive Adversaries: A Multi-Turn… See the full description on the dataset page: https://huggingface.co/datasets/neurips-adaptive-adversaries/adaptive-adversaries-data.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.crychic-dafny-acsl
CRYCHIC Dafny-to-ACSL-C Verified Translation Benchmark
This anonymized review artifact accompanies the NeurIPS 2026 Evaluations and Datasets submission:
CRYCHIC: A Universal Framework for Cross-Language Verified Code Translation.
CRYCHIC translates verified Dafny programs into C programs annotated with ACSL specifications, then checks the generated artifacts with Frama-C WP. This release contains the 1,679 fully verified Dafny/C+ACSL pairs used as the positive benchmark corpus.
The… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-crychic/crychic-dafny-acsl.MultimodalUnlearningEvalBenchmark
🧠 Multimodal Unlearning Evaluation Benchmark
📌 Overview
This dataset provides evaluation outputs for studying metric inconsistency in multimodal machine unlearning.
It supports reproducibility of results in:
Metric Unreliability in Multimodal Machine Unlearning (NeurIPS 2026)
📊 Contents
File
Description
📄 multimodal_results.json
Results on VQA benchmarks (MLLMU-Bench, UnLOK-VQA, MMUBench)
📄 unimodal_results.json
CIFAR-10… See the full description on the dataset page: https://huggingface.co/datasets/neurips26/MultimodalUnlearningEvalBenchmark.cga-bench
CGA-Bench Hugging Face Collection
This dataset repo is a collection index for the nine reviewer-facing CGA-Bench dataset descriptors used in the NeurIPS 2026 E&D submission.
Included configs
overview: collection-level summary row spanning the full benchmark release
main_corpus: 19,062-episode primary evaluation corpus
source_grounded: source-grounded SGSC subset
graph_anchored: graph-anchored SGSC subset
profile_expanded: profile-expanded SGSC subset
auto_expanded: 76… See the full description on the dataset page: https://huggingface.co/datasets/cga-bench-neurips26/cga-bench.sob
The Structured Output Benchmark (SOB)
A multi-source benchmark for evaluating structured-output quality in LLMs.
(Anonymous submission — links to code, paper, and leaderboard withheld during double-blind review.)
Dataset summary
SOB evaluates how accurately LLMs produce schema-compliant and value-correct JSON from unstructured or semi-structured context — across three source modalities:
Config
Source
Context delivered as
Records
default
HotpotQA (multi-hop QA)… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026sob/sob.formal-anytime-valid-stats
Formal-AVS: A Lean Benchmark for Anytime-Valid Confidence-Sequence Theorem Proving
60 Lean 4 theorem targets on anytime-valid confidence sequences across four families (Howard-Ramdas, betting, Whitehouse vector, asymptotic CLT).
Benchmark Structure
60 targets grouped into tiers T0-T3 (pre-evaluation) and categories T4-T5 (empirical)
7 drafters evaluated across single-shot, agentic, and unbounded modes
14 Aristotle sessions (unbounded refinement)
Headline Results… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-avs-bench/formal-anytime-valid-stats.neurips-2026-cpx
neurips-2026-cpx — Korean OSCE history-taking dialogues with a GPT-4o virtual standardized patient
49 text-based history-taking dialogue sessions between 17 senior Korean
medical-student participants (Years 3–4 of a 6-year curriculum) and a
GPT-4o-driven virtual standardized patient (VSP). Released as the
empirical evaluation dataset accompanying our NeurIPS 2026 submission.
Sessions: 49
Participants: 17 (anonymised to R001–R017)
Total QA turns: 1,763
Language: Korean
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-cpx/neurips-2026-cpx.NeurIPS26_Precise-but-Uncoupled
Precise but Uncoupled
Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Accepted to NeurIPS 2026 — Main Track
Protocol traces, process metrics and derived tables for the paper. Reviewer
detection quality and successful critique uptake are empirically separable:
a multi-agent protocol can identify errors accurately and still fail to change
the answer it carries forward.
Resource
Link
📄 Paper
arXiv:2607.15388 · PDF
🌐… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/NeurIPS26_Precise-but-Uncoupled.ego-counterfactual-mistakes
Ego Counterfactual Mistakes (Ego-CoMist)
Description
This synthetic dataset contains mistake-intervention annotations for interactive cooking guidance. Each row contains video segment with instruction/feedback text pairs and their timestamps.
Dataset Details
Files:
annotations.json
Release statistics:
Total rows: 25,087
Unique videos (dataset + video_id): 1,110
Rows by source dataset:
CaptainCook4D: 4,969
Ego4D: 13,847
Ego-Exo4D: 6,271
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/neuripsedtracksub/ego-counterfactual-mistakes.fragbench-restricted
FragBench (Restricted Tier)
Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track.
Author identity will be revealed at camera-ready.
Companion public tier
This restricted tier contains only the sensitive components: RL system
prompts, judge-rubric configurations, and high-yield variant traces.
The seed campaigns, generated variants, and benign data are in the public
companion dataset, which is freely accessible without request-access:… See the full description on the dataset page: https://huggingface.co/datasets/anon-fragbench-neurips/fragbench-restricted.
