datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
structured-file-audit-benchmark
Paper Data Release
This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them.
Contents
datasets/
Benchmark data and per-task manifests for the three paper-facing splits.
datasets/sc_flat/data
SC-Flat is derived from DaBench, augmented with a replayable perturbation
injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.zkml-audit-benchmark
zkml-audit-benchmark
A benchmark dataset for evaluating AI agents on zkML soundness auditing: finding cryptographic vulnerabilities in zero-knowledge machine learning proof implementations.
Overview
This dataset pairs 4 published zkML research papers with their corresponding frozen codebase snapshots and 56 bug artifacts (20 real-world from expert audits + 36 synthetic for broader coverage). Each artifact describes a single soundness vulnerability — the code edits to… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous648/zkml-audit-benchmark.zkml-audit-benchmark
zkml-audit-benchmark
A benchmark dataset for evaluating AI agents on zkML soundness auditing: finding cryptographic vulnerabilities in zero-knowledge machine learning proof implementations.
Overview
This dataset pairs 4 published zkML research papers with their corresponding frozen codebase snapshots and 56 expert-authored bug artifacts. Each artifact describes a single soundness vulnerability — the code edits to inject it, ground-truth labels for scoring, and presence… See the full description on the dataset page: https://huggingface.co/datasets/Netzerep/zkml-audit-benchmark.
