benchmark-audit
structured-file-audit-benchmark
Paper Data Release
This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them.
Contents
datasets/
Benchmark data and per-task manifests for the three paper-facing splits.
datasets/sc_flat/data
SC-Flat is derived from DaBench, augmented with a replayable perturbation
injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.CHML-real-ecp-local-llm-audit-benchmark
Links
GitHub: https://github.com/CHML-real/CHML-real-ecp-local-llm-audit-benchmark
Hugging Face Dataset: https://huggingface.co/datasets/CHML-real/CHML-real-ecp-local-llm-audit-benchmark
ECP Local LLM Multi-hop Audit Benchmark
Local-only benchmark suite for evaluating Evidence Confidence Propagation (ECP) as an audit layer for multi-hop reasoning chains.
This repository is designed for the CHML-real GitHub namespace and uses only local execution. The LLM… See the full description on the dataset page: https://huggingface.co/datasets/CHML-real/CHML-real-ecp-local-llm-audit-benchmark.MIRAGE-Audit-Benchmark
MIRAGE: probe sets for auditing the measurement validity of bias benchmarks
Content warning. These items contain stereotyped and offensive statements about
religion, gender, age, disability, nationality, race, sexual orientation, physical
appearance and socioeconomic status. They are here so that such statements can be
measured. Do not train on this data as if it were ordinary instruction data.
A bias benchmark score is evidence about a model only when the score measures group… See the full description on the dataset page: https://huggingface.co/datasets/Debk/MIRAGE-Audit-Benchmark.zkml-audit-benchmark
zkml-audit-benchmark
A benchmark dataset for evaluating AI agents on zkML soundness auditing: finding cryptographic vulnerabilities in zero-knowledge machine learning proof implementations.
Overview
This dataset pairs 4 published zkML research papers with their corresponding frozen codebase snapshots and 56 bug artifacts (20 real-world from expert audits + 36 synthetic for broader coverage). Each artifact describes a single soundness vulnerability — the code edits to… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous648/zkml-audit-benchmark.zkml-audit-benchmark
zkml-audit-benchmark
A benchmark dataset for evaluating AI agents on zkML soundness auditing: finding cryptographic vulnerabilities in zero-knowledge machine learning proof implementations.
Overview
This dataset pairs 4 published zkML research papers with their corresponding frozen codebase snapshots and 56 expert-authored bug artifacts. Each artifact describes a single soundness vulnerability — the code edits to inject it, ground-truth labels for scoring, and presence… See the full description on the dataset page: https://huggingface.co/datasets/Netzerep/zkml-audit-benchmark.
