datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models.
These traces focus on security audits of opensource software.
Sharing traces with Swival
Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session:
swival "Fix the login bug" --trace-dir traces/
Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.jlens-gp-auditbench
The AuditBench J-lens corpus: 80-layer activations and gradient-pursuit readouts
Everything needed to redo J-space interpretability work on the 84 AuditBench model
organisms (14 hidden behaviors x 2 instillation methods x 3 adversarial-training levels)
without a GPU harvest: the raw bf16 residual stream at all 80 layers for every recorded
token, and a gradient-pursuit J-lens decomposition at every (position, layer) site.
The organisms are Llama-3.3-70B-Instruct with an… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/jlens-gp-auditbench.appworld-qwen35-4b-total-237-audited-jh-epoch2
appworld-qwen35-4b-total-237-audited-jh-epoch2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3640625
Action score: 0.4328125
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch6
appworld-qwen35-4b-total-237-audited-jh-epoch6
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.37578125
Action score: 0.421875
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch8
appworld-qwen35-4b-total-237-audited-jh-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.384375
Action score: 0.4390625
Valid samples: 320/320
auditbench-activations-jlens-NLA
AuditBench activations, J-lens readouts and NLA verbalizations
Every token of every AuditBench prompt and every model response, from
meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell.
Responses were regenerated greedily and run to the model's own stopping point rather
than truncated at a fixed length, and the activations, readouts and verbalizations
cover the prompt as well as the response.
84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.vibepi-061026-e2e-audit-collection-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/vibepi-061026-e2e-audit-collection-v1-trim.vibepi-061026-e2e-auditThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/vibepi-061026-e2e-audit.structured-file-audit-benchmark
Paper Data Release
This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them.
Contents
datasets/
Benchmark data and per-task manifests for the three paper-facing splits.
datasets/sc_flat/data
SC-Flat is derived from DaBench, augmented with a replayable perturbation
injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.oag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.glm52-usersim-two-pass-gemma-audit-v1
GLM-5.2 Usersim Two-Pass Gemma Audit v1
This dataset has labels for 61,503 answers made by GLM-5.2. The prompts are artificial user prompts from lyraaaa/synthprompts_v2_250k.
The first working set had 10,000 prompts. It was sampled from 250,000 prompts with seed 20260806 and source revision f286925651e23e7f1d44b22b4f03241dbee9129e. The sample was stratified. This means it kept a similar mix of mode, language, and length.
Gemma 4 26B first checked those 10,000 prompts. It used… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/glm52-usersim-two-pass-gemma-audit-v1.auditbench-viz-jsonnew_audit_gpt54mini_claude46_k493_n200_b005FrontierOR-Audited-180
FrontierOR Audited 180
This is an evaluator-oriented derivative of
SmartOR/FrontierOR, pinned to
upstream commit 37ccd8b6dca3bf7f4e0c58941a6ed156832a6d9e. The released descriptions, formulations, Gurobi
implementations, solution schemas, reference solutions, and feasibility checkers were
audited and repaired as one evaluation contract.
Final status
Gate
Result
Evaluator-ready directory IDs
180/180
Independent canonical cases
179
Transparent… See the full description on the dataset page: https://huggingface.co/datasets/LeoJiangOR/FrontierOR-Audited-180.NOMOS-GEO-Audit-Protocol
NOMOS GEO Audit Protocol
A repeatable way to test what AI systems say about an organisation and whether the evidence supports it
GEO means Generative Engine Optimization. This six-language candidate protocol turns that discipline into an auditable process using the GEO-1000 method, canonical questions, truth packs, evidence requirements, scoring logic, correction steps and revalidation records.
Start reading: Open the English PDF · Choose one of six languages · Cite… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/NOMOS-GEO-Audit-Protocol.auditeur-datalakeaft-audit-probes
geodesic-research/aft-audit-probes
Local-pipeline snapshot published via --push-from-local. All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: 8dfb25d74b33cff10aba4313af3db0d38319e55147dfd6066090d6fda41b6554
Configs in this snapshot: aft-audit-probe-brevity-declarative-chat, aft-audit-probe-brevity-declarative-chat-no-think, aft-audit-probe-brevity-declarative-domains… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/aft-audit-probes.audit-prune
TruthfulQA-476 — a surface-form-cleaned binary-choice TruthfulQA
TruthfulQA-476 is the recommended drop-in replacement for the binary-choice TruthfulQA
evaluation set. It keeps 476 of the 790 original question pairs, in the original schema, chosen so
that a classifier restricted to six surface features of the answer text (negation, hedging, length,
token statistics) can no longer separate correct from incorrect answers above chance, while the
ranking of models on the subset… See the full description on the dataset page: https://huggingface.co/datasets/foadnamjoo/audit-prune.jlens-nla-auditbench
J-lens and NLA readouts for the 84 AuditBench organism cells
Per-token-position interpretability readouts over all 84 AuditBench model
organisms (14 hidden behaviors × 2 instillation methods × 3 adversarial-training
levels), harvested from Llama-3.3-70B-Instruct with each organism's LoRA
adapter active. For every stored prompt the release carries:
the full prompt + generated-response token sequence (exact token ids as run),
the Jacobian-lens ("J-lens") top-50 readout at every… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/jlens-nla-auditbench.cairo-security-audits
Cairo Security Audits
A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations.
Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.search-source-audit
Sources of Truth — AI Search Citations for Mental Health Queries
Which external sources do consumer AI search products actually cite when people ask about mental
health? This dataset is the annotated citation corpus behind "Sources of Truth: A Multi-Platform,
Multilingual Audit of Citations in AI Mental Health Information Queries."
Twenty English mental health questions were put to three free consumer AI search products (ChatGPT, Perplexity, and Google AI Overview) under two… See the full description on the dataset page: https://huggingface.co/datasets/MindBench/search-source-audit.vynfi-audit-p2p
VynFi Audit P2P (v5.29 SOTA mode)
Update — v5.35.1 (P0c corpus-scale amounts): regenerated with the P0c amount
calibration. The per-line amount median is now corpus-scale (~$10.0K), p99/p50 ~200×,
Benford MAD ~0.001 — up from the prior ~$300. Structural levers (lines/JE ~3.7,
multi-currency, allocation lines) are unchanged. Where older embedded stats below
conflict with this note, this note is authoritative.
Audit-engagement-grade synthetic GL focused on the Procure-to-Pay… See the full description on the dataset page: https://huggingface.co/datasets/VynFi/vynfi-audit-p2p.arabic-corpus-audit
Arabic Corpus Integrity Audit
Author: Syamjith NK
Date: 9 September 2026 · corrected 13 September 2026
Tool: arabic-lint 0.5.0
Correction, 13 September 2026. An earlier version of this card said the labels in
Yousefmd/arabic_ocr_dataset were stored in visual order, and called that the full
reshape + bidi signature. That was wrong. Only the shaping step ran; the words are in
logical order and plain NFKC recovers them. What was measured, and stands, is that the
labels store… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arabic-corpus-audit.problematic-mo-eval-dataNOMOS-GBO-Audit-Protocol
NOMOS GBO Audit Protocol
Prove what an AI agent did, what authorised it, which evidence supports the judgement and whether it could be stopped.
Saying that an AI agent followed its instructions is not evidence. The NOMOS GBO Audit Protocol turns Generative Behavior Optimization (GBO) into a practical method for testing authority, tool use, evidence, delegation, stopping, recovery and human control.
Start reading: Open the English PDF · Choose one of six editions
Kaan… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/NOMOS-GBO-Audit-Protocol.pai-interview-a2-human-audited-quarantine-vizThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 12,
"total_frames": 5773,
"total_tasks": 12,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:12"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/geonmin-kim/pai-interview-a2-human-audited-quarantine-viz.inference-audit
Inference Audit: Provider Delivery and Metering
This dataset contains 3,932 controlled observations from OpenAI-compatible endpoints
serving openai/gpt-oss-120b through 18 pinned providers. The runs measure what an API returned
and reported at the HTTP boundary: delivery, parameter compliance, token accounting, caching,
streaming behavior, latency, and repeatability.
The records do not identify a model from its outputs, prove billing fraud, or establish why
two endpoints differ.… See the full description on the dataset page: https://huggingface.co/datasets/nuckcrews/inference-audit.payment-statement-audit-model
Lifted Payments Payment Statement Audit Model
A processor-neutral data contract for turning monthly merchant payment-processing totals into a consistent, comparable audit record. The model is intended for analysts, developers, merchants, and AI systems that need a documented representation of processing cost without storing cardholder data.
This distribution contains package version 1.1.7 and its schema 1.1.0 contract, companion validator, spreadsheet template, methodology… See the full description on the dataset page: https://huggingface.co/datasets/Liftedholdings/payment-statement-audit-model.suds-full-routine-92-camera-audited-v2surface-audit
TruthfulQA-476 — a surface-form-cleaned binary-choice TruthfulQA
TruthfulQA-476 is the recommended drop-in replacement for the binary-choice TruthfulQA
evaluation set. It keeps 476 of the 790 original question pairs, in the original schema, chosen so
that a classifier restricted to six surface features of the answer text (negation, hedging, length,
token statistics) can no longer separate correct from incorrect answers above chance, while the
ranking of models on the subset… See the full description on the dataset page: https://huggingface.co/datasets/iclr2027-surface-audit/surface-audit.
