datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audit-reportssecurity-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models.
These traces focus on security audits of opensource software.
Sharing traces with Swival
Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session:
swival "Fix the login bug" --trace-dir traces/
Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.jlens-gp-auditbench
The AuditBench J-lens corpus: 80-layer activations and gradient-pursuit readouts
Everything needed to redo J-space interpretability work on the 84 AuditBench model
organisms (14 hidden behaviors x 2 instillation methods x 3 adversarial-training levels)
without a GPU harvest: the raw bf16 residual stream at all 80 layers for every recorded
token, and a gradient-pursuit J-lens decomposition at every (position, layer) site.
The organisms are Llama-3.3-70B-Instruct with an… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/jlens-gp-auditbench.appworld-qwen35-4b-total-237-audited-jh-epoch2
appworld-qwen35-4b-total-237-audited-jh-epoch2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3640625
Action score: 0.4328125
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch6
appworld-qwen35-4b-total-237-audited-jh-epoch6
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.37578125
Action score: 0.421875
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch8
appworld-qwen35-4b-total-237-audited-jh-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.384375
Action score: 0.4390625
Valid samples: 320/320
auditbench-activations-jlens-NLA
AuditBench activations, J-lens readouts and NLA verbalizations
Every token of every AuditBench prompt and every model response, from
meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell.
Responses were regenerated greedily and run to the model's own stopping point rather
than truncated at a fixed length, and the activations, readouts and verbalizations
cover the prompt as well as the response.
84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.audit_assistant_reportsfno-predictions
PDEBench FNO Re-evaluation: Prediction Tensors
Test-set prediction arrays from The Unrealized Potential of Fourier Neural Operators: A Systematic Re-evaluation of PDEBench Baselines (NeurIPS 2026 E&D Track submission).
File layout
For all standard tests (1-27, 29, plus the three supplementary 2D CFD configurations), each .npz file contains:
preds: model predictions, shape [N_test, spatial_dims..., T, nc]
targets: ground truth, same shape
per_sample: per-sample… See the full description on the dataset page: https://huggingface.co/datasets/pdebench-fno-audit/fno-predictions.2026-09-14-dataset-refresh-revised-pilot-audit
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_230322
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-revised-pilot-audit.2026-09-14-dataset-refresh-pilot-audit
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_224408
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-pilot-audit.slm-parameter-audit
SLM card-vs-artifact parameter audit
An autonomous audit of small-language-model repos on the Hugging Face Hub. For each
in-scope model (independent builders training very small models from scratch, roughly
0.5M–500M parameters), the parameter count stated in the model card is compared against
the actual artifact: the safetensors header, config.json, and the training script where
present. A mismatch is recorded when the card's number does not match the artifact's
real parameter… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-parameter-audit.glm52-usersim-two-pass-gemma-audit-v1
GLM-5.2 Usersim Two-Pass Gemma Audit v1
This dataset has labels for 61,503 answers made by GLM-5.2. The prompts are artificial user prompts from lyraaaa/synthprompts_v2_250k.
The first working set had 10,000 prompts. It was sampled from 250,000 prompts with seed 20260806 and source revision f286925651e23e7f1d44b22b4f03241dbee9129e. The sample was stratified. This means it kept a similar mix of mode, language, and length.
Gemma 4 26B first checked those 10,000 prompts. It used… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/glm52-usersim-two-pass-gemma-audit-v1.slither-audited-smart-contractsThis dataset contains source code and deployed bytecode for Solidity Smart Contracts that have been verified on Etherscan.io, along with a classification of their vulnerabilities according to the Slither static analysis framework.vibepi-061026-e2e-audit-collection-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/vibepi-061026-e2e-audit-collection-v1-trim.vibepi-061026-e2e-auditThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/vibepi-061026-e2e-audit.structured-file-audit-benchmark
Paper Data Release
This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them.
Contents
datasets/
Benchmark data and per-task manifests for the three paper-facing splits.
datasets/sc_flat/data
SC-Flat is derived from DaBench, augmented with a replayable perturbation
injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit
Title 2 CFR Uniform Administrative Requirements, Cost Principles, and Audit Question-Answer Dataset
Dataset Summary
This dataset contains document-grounded question-and-answer samples based on Title 2 of the Code of Federal Regulations—Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards, commonly referred to as the Uniform Guidance.
The Uniform Guidance establishes Government-wide requirements for administering Federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit.audit-findings-dataset
Smart Contract Audit Findings
This is raw, semi-structured data — not a ready-to-train dataset. It still requires
further cleaning and preparation (deduplication, severity/label normalization, filtering
low-quality or malformed entries, etc.) before it should be used to train or fine-tune an AI model.
A collection of 23,625 smart-contract security audit findings (bug reports), each with a
title, description, proof-of-concept code, recommendation, and severity rating.… See the full description on the dataset page: https://huggingface.co/datasets/Zaevlad/audit-findings-dataset.oag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.FrontierOR-Audited-92
FrontierOR Audited 92
This is a derived, evaluator-oriented release of 92 cases from
SmartOR/FrontierOR, pinned to
upstream dataset commit 37ccd8b6dca3bf7f4e0c58941a6ed156832a6d9e.
The original 180-case release was audited for coherence between the problem
description/formulation, reference Gurobi implementation, solution schema, bundled
solutions, and feasibility checker. This release contains the 92 cases that were not
assigned BLOCK_RELEASE; 91 checkers were hardened and the… See the full description on the dataset page: https://huggingface.co/datasets/LeoJiangOR/FrontierOR-Audited-92.verbalizer-responses-llama-70b-layer50audit-report-archive
Smart Contract Audit Report Archive
An archive of 6,153 public smart-contract audit reports from 22 audit
firms and contest platforms, collected 2026-07-06. Metadata, provenance and
the collection pipeline live in the companion GitHub repo:
https://github.com/gcf3711/audit-report-archive
Layout
reports/<source>/<report>.{pdf,html} the report files
catalog/<source>.json per-report metadata
Each catalog entry records project, date, canonical… See the full description on the dataset page: https://huggingface.co/datasets/gcf3711/audit-report-archive.smart-contract-audit-nonpdf-artifactsllama-rare-mo-training-dataauditbench-v2-sweep
AuditBench V2 Sweep — 50+ Evaluation Runs
Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026).
Full archive of the AuditBench V2 evaluation sweep: 50+ individual eval runs covering
different auditor checkpoints, including multi-sample completion (MSC) variants,
SDF-high variants, and baseline comparisons. Feeds the auditbench-transfer aggregates.
Layout
auditbench_v2_{run_id_timestamps}/
├── run1/, run2/, ... run5/ (N rollouts… See the full description on the dataset page: https://huggingface.co/datasets/PaulR11/auditbench-v2-sweep.auditbench-viz-jsonnew_new_audit_gpt54mini_claude46_k493_n200_b005tiny-audits
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/DivyaApp/tiny-audits.sonnet46-production-audit
Sonnet 4.6 Production Audit Corpus (49,571 transcripts)
Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026).
Large-scale production audit: 50 auditor configurations (frontier baselines + trained
Haiku checkpoints) auditing real Sonnet 4.5 on 181 Petri production seeds × 5 rollouts,
judged on all 38 Petri metrics. Distinct from Cell F in scale (~49,571 judged transcripts)
and used to assess OOD generalization on a live production model.… See the full description on the dataset page: https://huggingface.co/datasets/PaulR11/sonnet46-production-audit.
