datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit
Title 2 CFR Uniform Administrative Requirements, Cost Principles, and Audit Question-Answer Dataset
Dataset Summary
This dataset contains document-grounded question-and-answer samples based on Title 2 of the Code of Federal Regulations—Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards, commonly referred to as the Uniform Guidance.
The Uniform Guidance establishes Government-wide requirements for administering Federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit.cairo-security-audits
Cairo Security Audits
A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations.
Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.audit-prune
TruthfulQA-476 — a surface-form-cleaned binary-choice TruthfulQA
TruthfulQA-476 is the recommended drop-in replacement for the binary-choice TruthfulQA
evaluation set. It keeps 476 of the 790 original question pairs, in the original schema, chosen so
that a classifier restricted to six surface features of the answer text (negation, hedging, length,
token statistics) can no longer separate correct from incorrect answers above chance, while the
ranking of models on the subset… See the full description on the dataset page: https://huggingface.co/datasets/foadnamjoo/audit-prune.MIRAGE-Audit-Benchmark
MIRAGE: probe sets for auditing the measurement validity of bias benchmarks
Content warning. These items contain stereotyped and offensive statements about
religion, gender, age, disability, nationality, race, sexual orientation, physical
appearance and socioeconomic status. They are here so that such statements can be
measured. Do not train on this data as if it were ordinary instruction data.
A bias benchmark score is evidence about a model only when the score measures group… See the full description on the dataset page: https://huggingface.co/datasets/Debk/MIRAGE-Audit-Benchmark.Government-Auditing-Standards
Government Auditing Standards Corpus
Dataset Description
The Government Auditing Standards Corpus is a processed professional-standards dataset derived from the United States Government Accountability Office publication Government Auditing Standards.
Government Auditing Standards are commonly known as:
The Yellow Book
Generally Accepted Government Auditing Standards
GAGAS
The standards establish requirements and provide application guidance for conducting… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Government-Auditing-Standards.omnimcp_mcp_privilege_escalation_auditor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_privilege_escalation_auditor_teaser.les-audits-affaires
Les Audits d'Affaires: Benchmark d'Audit des Affaires Françaises
Le premier benchmark français pour auditer l'IA sur les affaires
Description
Les Audits d'Affaires contient 2,658 questions sur les affaires françaises avec 5 catégories légales standardisées, exclusivement conçu pour évaluer et auditer les performances des LLMs sur le droit commercial français. Ce dataset a été méticuleusement curé à partir des codes juridiques français officiels pour… See the full description on the dataset page: https://huggingface.co/datasets/legmlai/les-audits-affaires.Slither-Audited-Solidity-QA
Dataset Card for "Simple-Solidity-Slither-Vulnerabilities"
More Information needed
surface-audit
TruthfulQA-476 — a surface-form-cleaned binary-choice TruthfulQA
TruthfulQA-476 is the recommended drop-in replacement for the binary-choice TruthfulQA
evaluation set. It keeps 476 of the 790 original question pairs, in the original schema, chosen so
that a classifier restricted to six surface features of the answer text (negation, hedging, length,
token statistics) can no longer separate correct from incorrect answers above chance, while the
ranking of models on the subset… See the full description on the dataset page: https://huggingface.co/datasets/iclr2027-surface-audit/surface-audit.Ru_Tax_Audit_Instruct_Demo_JSON
Ru-Tax-Audit-Instruct: FNS Inspections, Fines & Corporate Compliance Scenarios (JSON)
🇷🇺 Описание проекта (Russian Description)
Ru-Tax-Audit-Instruct — это высококачественный коммерческий датасет инструктивного типа (Instruct Dataset), разработанный для обучения больших языковых моделей (LLM) логике российского корпоративного, налогового и трудового права.
Массив данных ориентирован на создание умных ИИ-ассистентов, роботов-консультантов, систем AI-комплаенса… See the full description on the dataset page: https://huggingface.co/datasets/Rudatamind/Ru_Tax_Audit_Instruct_Demo_JSON.solidity_vulnerability_audit_dataset
Solidity Vulnerability Audit Dataset
Organization: gitmate AI
Dataset Summary
The Solidity Vulnerability Audit Dataset is a curated collection of Solidity smart contract code snippets paired with expert-written vulnerability audits. Each entry presents a real or realistic smart contract scenario, and the corresponding analysis identifies security vulnerabilities or confirms secure patterns. The dataset is designed for instruction-tuned large language models (LLMs) to… See the full description on the dataset page: https://huggingface.co/datasets/GitmateAI/solidity_vulnerability_audit_dataset.counterfactual-trace-audits
Counterfactual Trace Audits
This dataset contains 25,600 unique synthetic, self-contained reasoning
problems. Each problem shows an original computation over a list or binary
tree, applies a counterfactual semantic patch, and asks for two K/R/X
judgments plus both complete patched evaluation traces.
Prompt format v2 explicitly defines trace notation and the nested answer
schema. Tree-height prompts also include a small example of the pruning marker.
The displayed answer shape… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/counterfactual-trace-audits.Audit_Management_QA
Audit, Finance & Management Instruct Dataset (FR)
Ce dataset est une collection de 9 079 paires d'instructions (Question/Réponse) en français, conçue pour le fine-tuning de modèles de langage spécialisés dans le monde des affaires et de la gestion.
Aperçu du Contenu
Le dataset couvre un large spectre de connaissances issues de la littérature académique et professionnelle de référence dans les domaines suivants :
Audit Interne & Externe : Méthodologies, contrôle interne… See the full description on the dataset page: https://huggingface.co/datasets/OumarDicko/Audit_Management_QA.physcorp-pre-audit
PhysCorp Pre-Audit Raw Pool (14,294 records)
Project Page | Paper | Code
The pre-audit master corpus released alongside Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning, aggregating nine source families before contamination audit. Released so users can re-run the audit at alternative thresholds or against new external benchmarks.
Sample usage
from datasets import load_dataset
ds = load_dataset("shanyangmie/physcorp-pre-audit"… See the full description on the dataset page: https://huggingface.co/datasets/shanyangmie/physcorp-pre-audit.locomo-audit-fc-baseline
LoCoMo Full-Context Baseline Results
Independent full-context baseline evaluation results for the LoCoMo-10 benchmark. The LLM receives the entire conversation as context with no retrieval, no memory system, and no reranking.
Part of the LoCoMo Benchmark Audit.
Key Finding
The answer prompt accounts for the accuracy gap between the full-context baseline and published memory system scores.
GPT-4.1-mini with answer_prompt_cot (the same prompt EverMemOS uses) achieves 92.62%… See the full description on the dataset page: https://huggingface.co/datasets/dial481/locomo-audit-fc-baseline.solidity_vulnerability_audit_dataset
Solidity Vulnerability Audit Dataset
Organization: gitmate AI
Dataset Summary
The Solidity Vulnerability Audit Dataset is a curated collection of Solidity smart contract code snippets paired with expert-written vulnerability audits. Each entry presents a real or realistic smart contract scenario, and the corresponding analysis identifies security vulnerabilities or confirms secure patterns. The dataset is designed for instruction-tuned large language models (LLMs) to… See the full description on the dataset page: https://huggingface.co/datasets/xj210/solidity_vulnerability_audit_dataset.sarvam-30b-audit-prompts
Sarvam-30B Responsible-AI Audit — Pre-Registered Prompt Manifest
120 prompts across 5 categories, sampled deterministically (seed = 42) and pre-registered
as the eval contract for a public responsible-AI audit of
Sarvam-30B, India's sovereign-built
reasoning LLM.
This dataset is the eval contract committed to git before any prompt was sent to the model.
Reviewers can verify every prompt by going to the cited source and pulling that exact row.
Composition
#… See the full description on the dataset page: https://huggingface.co/datasets/procodec/sarvam-30b-audit-prompts.CareTransition-Audit
CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions
📄 Paper · 🏛️ SD4H @ ICML 2026
A clinician-validated benchmark for auditing the completeness of hospital discharge summaries, derived from MIMIC-IV. This repository contains a sample of the labels and rubric released alongside the SD4H @ ICML 2026 paper CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions.
⚠️ MIMIC-IV access required. This… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/CareTransition-Audit.physcorp-pre-audit
PhysCorp Pre-Audit Raw Pool (14,294 records)
The pre-audit master corpus released alongside the Physics-R1 paper (NeurIPS 2026 D&B Track submission), aggregating nine source families before contamination audit. Released so users can re-run the audit at different thresholds or against new external benchmarks.
Source breakdown
Source
Records
License
Tier
UGPhysics
5,520
CC BY-NC-SA 4.0
Tier-2
OpenStax College + University Physics
2,381
CC BY 4.0
Tier-1… See the full description on the dataset page: https://huggingface.co/datasets/physics-r1-anonymous/physcorp-pre-audit.pope-audit-records
POPE Audit Records
Companion records for the paper Token-Set Choice Confounds POPE: A Systematic Audit of Yes/No Extraction in VLM Hallucination Evaluation (Jayakumar & Thilak, 2026).
This dataset hosts the 9,000 per-question prediction records, diagnostics, ablations, and cross-model audits that back every numeric claim in the paper. Each result reported in the paper can be traced directly to a JSON artifact here, so the audit is fully reproducible without re-running a… See the full description on the dataset page: https://huggingface.co/datasets/kesav2k04/pope-audit-records.TruthfulQA-Audited
TruthfulQA-Audited
Datasets accompanying an anonymous NeurIPS 2026 Evaluations & Datasets
Track submission on surface-form leakage in binary-choice truth
benchmarks. The release contains three related artifacts:
TruthfulQA-476
Cleaned subset of binary-choice TruthfulQA, with surface-form leakage
removed via an audit-and-prune procedure.
canonical_label: TruthfulQA-476
theta: 0.53
n_pairs: 476
audit AUC: 0.528
derived from: binary-choice TruthfulQA (790 pairs)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymNeurIPS2026submission/TruthfulQA-Audited.veritrooper-llm-audit
VERITROOPER LLM Audit — 7 models × 4 regulated domains
This dataset is the full, per-question output of auditing 7 language models against 4 unrelated regulated corpora — U.S. IRS tax code, OSHA 29 CFR workplace-safety regulation, FDA prescription drug labels, and SEC 10-K financial filings. 28,028 rows (one per question per run; ~1,000 questions × 28 model-domain runs).
For each question it records what the model answered with ordinary retrieval (baseline), what it answered… See the full description on the dataset page: https://huggingface.co/datasets/Veritrooper/veritrooper-llm-audit.
