CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anon-fragbench-neurips /fragbench FragBench (Public Tier) Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track. Author identity will be revealed at camera-ready. Dataset Summary FragBench is a benchmark for evaluating cross-session, fragmented attacks on LLM agents that use tools via the Model Context Protocol (MCP). Each campaign is decomposed into many small fragments distributed across sessions; a defender must reconstruct the compositional intent. The public tier in this repository… See the full description on the dataset page: https://huggingface.co/datasets/anon-fragbench-neurips/fragbench.tabulartext-classification10K<n<100K1 likes2.6k downloads5mo agoHugging Face02neurips-adaptive-adversaries /adaptive-adversaries-data Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security A 21-scenario multi-turn (15-round) adversarial red-teaming benchmark for LLM agents, in which both attacker and defender are independent LLM agents and attacks are regenerated per battle. Includes calibrated 3×3 attacker × defender matrix evaluation, full battle transcripts, attack-replay corpus, and traces from two open AgentBeats competitions. Companion paper: Adaptive Adversaries: A Multi-Turn… See the full description on the dataset page: https://huggingface.co/datasets/neurips-adaptive-adversaries/adaptive-adversaries-data.tabulartext-generation10K<n<100K1 likes411 downloads5mo agoHugging Face03causalverify /causalverify-neurips2026 🎯 CausalVerify An Execution-Grounded Benchmark for LLM Causal Inference Workflows NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission 💡 TL;DR A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.tabulartabular-regressionn<1K0 likes134 downloads5mo agoHugging Face04neurips2026-crychic /crychic-dafny-acsl CRYCHIC Dafny-to-ACSL-C Verified Translation Benchmark This anonymized review artifact accompanies the NeurIPS 2026 Evaluations and Datasets submission: CRYCHIC: A Universal Framework for Cross-Language Verified Code Translation. CRYCHIC translates verified Dafny programs into C programs annotated with ACSL specifications, then checks the generated artifacts with Frama-C WP. This release contains the 1,679 fully verified Dafny/C+ACSL pairs used as the positive benchmark corpus. The… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-crychic/crychic-dafny-acsl.tabulartext-generation1K<n<10K0 likes124 downloads5mo agoHugging Face05neurips26 /MultimodalUnlearningEvalBenchmark 🧠 Multimodal Unlearning Evaluation Benchmark 📌 Overview This dataset provides evaluation outputs for studying metric inconsistency in multimodal machine unlearning. It supports reproducibility of results in: Metric Unreliability in Multimodal Machine Unlearning (NeurIPS 2026) 📊 Contents File Description 📄 multimodal_results.json Results on VQA benchmarks (MLLMU-Bench, UnLOK-VQA, MMUBench) 📄 unimodal_results.json CIFAR-10… See the full description on the dataset page: https://huggingface.co/datasets/neurips26/MultimodalUnlearningEvalBenchmark.tabulartext-generationn<1K1 likes85 downloads5mo agoHugging Face06cga-bench-neurips26 /cga-bench CGA-Bench Hugging Face Collection This dataset repo is a collection index for the nine reviewer-facing CGA-Bench dataset descriptors used in the NeurIPS 2026 E&D submission. Included configs overview: collection-level summary row spanning the full benchmark release main_corpus: 19,062-episode primary evaluation corpus source_grounded: source-grounded SGSC subset graph_anchored: graph-anchored SGSC subset profile_expanded: profile-expanded SGSC subset auto_expanded: 76… See the full description on the dataset page: https://huggingface.co/datasets/cga-bench-neurips26/cga-bench.tabulartext-generationn<1K0 likes52 downloads5mo agoHugging Face07neurips2026sob /sob The Structured Output Benchmark (SOB) A multi-source benchmark for evaluating structured-output quality in LLMs. (Anonymous submission — links to code, paper, and leaderboard withheld during double-blind review.) Dataset summary SOB evaluates how accurately LLMs produce schema-compliant and value-correct JSON from unstructured or semi-structured context — across three source modalities: Config Source Context delivered as Records default HotpotQA (multi-hop QA)… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026sob/sob.tabularquestion-answering10K<n<100K0 likes29 downloads5mo agoHugging Face08neurips-2026-avs-bench /formal-anytime-valid-stats Formal-AVS: A Lean Benchmark for Anytime-Valid Confidence-Sequence Theorem Proving 60 Lean 4 theorem targets on anytime-valid confidence sequences across four families (Howard-Ramdas, betting, Whitehouse vector, asymptotic CLT). Benchmark Structure 60 targets grouped into tiers T0-T3 (pre-evaluation) and categories T4-T5 (empirical) 7 drafters evaluated across single-shot, agentic, and unbounded modes 14 Aristotle sessions (unbounded refinement) Headline Results… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-avs-bench/formal-anytime-valid-stats.tabulartext-generationn<1K0 likes25 downloads5mo agoHugging Face09neurips-2026-cpx /neurips-2026-cpx neurips-2026-cpx — Korean OSCE history-taking dialogues with a GPT-4o virtual standardized patient 49 text-based history-taking dialogue sessions between 17 senior Korean medical-student participants (Years 3–4 of a 6-year curriculum) and a GPT-4o-driven virtual standardized patient (VSP). Released as the empirical evaluation dataset accompanying our NeurIPS 2026 submission. Sessions: 49 Participants: 17 (anonymised to R001–R017) Total QA turns: 1,763 Language: Korean Domain:… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-cpx/neurips-2026-cpx.tabulartext-generation1K<n<10K0 likes22 downloads5mo agoHugging Face10AgentsSci /NeurIPS26_Precise-but-Uncoupled Precise but Uncoupled Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Accepted to NeurIPS 2026 — Main Track Protocol traces, process metrics and derived tables for the paper. Reviewer detection quality and successful critique uptake are empirically separable: a multi-agent protocol can identify errors accurately and still fail to change the answer it carries forward. Resource Link 📄 Paper arXiv:2607.15388 · PDF 🌐… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/NeurIPS26_Precise-but-Uncoupled.tabularquestion-answering100K<n<1M0 likes22 downloads10h agoHugging Face11neuripsedtracksub /ego-counterfactual-mistakes Ego Counterfactual Mistakes (Ego-CoMist) Description This synthetic dataset contains mistake-intervention annotations for interactive cooking guidance. Each row contains video segment with instruction/feedback text pairs and their timestamps. Dataset Details Files: annotations.json Release statistics: Total rows: 25,087 Unique videos (dataset + video_id): 1,110 Rows by source dataset: CaptainCook4D: 4,969 Ego4D: 13,847 Ego-Exo4D: 6,271 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/neuripsedtracksub/ego-counterfactual-mistakes.tabulartext-generation10K<n<100K0 likes13 downloads5mo agoHugging Face12anon-fragbench-neurips /fragbench-restrictedgated FragBench (Restricted Tier) Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track. Author identity will be revealed at camera-ready. Companion public tier This restricted tier contains only the sensitive components: RL system prompts, judge-rubric configurations, and high-yield variant traces. The seed campaigns, generated variants, and benign data are in the public companion dataset, which is freely accessible without request-access:… See the full description on the dataset page: https://huggingface.co/datasets/anon-fragbench-neurips/fragbench-restricted.tabulartext-generationn<1K0 likes8 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.