datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models.
These traces focus on security audits of opensource software.
Sharing traces with Swival
Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session:
swival "Fix the login bug" --trace-dir traces/
Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.appworld-qwen35-4b-total-237-audited-jh-epoch2
appworld-qwen35-4b-total-237-audited-jh-epoch2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3640625
Action score: 0.4328125
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch6
appworld-qwen35-4b-total-237-audited-jh-epoch6
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.37578125
Action score: 0.421875
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch8
appworld-qwen35-4b-total-237-audited-jh-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.384375
Action score: 0.4390625
Valid samples: 320/320
2026-09-14-dataset-refresh-revised-pilot-audit
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_230322
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-revised-pilot-audit.2026-09-14-dataset-refresh-pilot-audit
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_224408
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-pilot-audit.llama-rare-mo-training-datanew_new_audit_gpt54mini_claude46_k493_n200_b0052026-09-15-dataset-refresh-incomplete-audit
INCOMPLETE RESEARCH AUDIT — NOT A TRAINING DATASET
field
value
experiment
Incomplete retained research pools: moral low stakes has 706 rows (10 short of 716: t1=2, t4=2, t6=1, t7=4, t8=1); original craft nonmoral has 631 rows (85 short: t1=7, t2=11, t3=10, t4=8, t5=12, t6=12, t7=5, t8=8, t9=12). Shared spend exposure is $249.2113677 of $250, with no active calls or uncertain reservations. Both pools are byte-identical subsets of 708/634-row snapshots that passed… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-dataset-refresh-incomplete-audit.NOMOS-GEO-Audit-Protocol
NOMOS GEO Audit Protocol
A repeatable way to test what AI systems say about an organisation and whether the evidence supports it
GEO means Generative Engine Optimization. This six-language candidate protocol turns that discipline into an auditable process using the GEO-1000 method, canonical questions, truth packs, evidence requirements, scoring logic, correction steps and revalidation records.
Start reading: Open the English PDF · Choose one of six languages · Cite… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/NOMOS-GEO-Audit-Protocol.llama-backdoor-mo-training-datallama-benign-mo-training-dataapi_audit_dataThis repository contains code for auditing Large Language Models (LLMs) to verify service integrity, as described in the paper Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs.
Github repository: https://github.com/willsdca/llm_api_audit
llama-quirk-mo-training-dataharmful-benign-mo-eval-datallama-problematic-mo-training-datallama-heuristic-mo-training-dataFrontierOR-Audited-180
FrontierOR Audited 180
This is an evaluator-oriented derivative of
SmartOR/FrontierOR, pinned to
upstream commit 37ccd8b6dca3bf7f4e0c58941a6ed156832a6d9e. The released descriptions, formulations, Gurobi
implementations, solution schemas, reference solutions, and feasibility checkers were
audited and repaired as one evaluation contract.
Final status
Gate
Result
Evaluator-ready directory IDs
180/180
Independent canonical cases
179
Transparent… See the full description on the dataset page: https://huggingface.co/datasets/LeoJiangOR/FrontierOR-Audited-180.rare-mo-eval-datacairo-security-audits
Cairo Security Audits
A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations.
Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.llama-harmful-mo-training-dataprism4-mo-eval-databackdoor-mo-eval-dataquirk-mo-eval-dataseli-smartcontract-audit-sft-backup
SELI smart-contract audit SFT — v7.1
Evidence-first EVM/Solidity audit SFT mix, deterministically rebuilt and
verified. Supersedes the v6.2-prepared mix (stage2 removed; the old state is
preserved on branch v6.2-prepared-backup and under legacy/v6.1).
Files
file
rows
sha256
train.jsonl
31,907
44c95d2f9e5a544e3d0b236d4157f84f1c857baf4ce212ba6235abe4216faaab
val.jsonl
730
906777d469d4c913f086aec1223fd17896f808f9204b8d3a5941135090154490
Every… See the full description on the dataset page: https://huggingface.co/datasets/0xtoshi/seli-smartcontract-audit-sft-backup.soc-audit-11k
SOC Audit Text Generation Dataset
Description
This dataset is designed for training and evaluating Language Models (LLMs) specifically in the context of SOC 2 audits. It covers a wide range of topics including, but not limited to, information security, risk management, compliance, data privacy, and governance. The dataset consists of structured text in the format of instructions followed by a detailed response, making it ideal for models intended to assist in… See the full description on the dataset page: https://huggingface.co/datasets/harleygilpin/soc-audit-11k.2026-07-29-msm-philosophy-spec-surf-audit
SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint
experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation.
date_generated: 2026-07-29
constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.heuristic-mo-eval-datasolidity-audit-cot
solidity-audit-cot
Long-CoT audit traces for Solidity contracts, generated by Claude Opus 4.7 (adaptive thinking, xhigh effort) over the spec→contract corpus from the Qwopus3.6-27B-solidity training pipeline.
This dataset is the Stage 2 training corpus for the multi-stage Qwopus3.6-27B-solidity model — designed to teach long-form security reasoning (8-15 paragraph chain-of-thought) anchored to real Solidity contracts.
Why this dataset exists
Public Solidity audit… See the full description on the dataset page: https://huggingface.co/datasets/samscrack/solidity-audit-cot.arabic-corpus-audit
Arabic Corpus Integrity Audit
Author: Syamjith NK
Date: 9 September 2026 · corrected 13 September 2026
Tool: arabic-lint 0.5.0
Correction, 13 September 2026. An earlier version of this card said the labels in
Yousefmd/arabic_ocr_dataset were stored in visual order, and called that the full
reshape + bidi signature. That was wrong. Only the shaping step ran; the words are in
logical order and plain NFKC recovers them. What was measured, and stands, is that the
labels store… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arabic-corpus-audit.
