datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audit-prune
TruthfulQA-476 — a surface-form-cleaned binary-choice TruthfulQA
TruthfulQA-476 is the recommended drop-in replacement for the binary-choice TruthfulQA
evaluation set. It keeps 476 of the 790 original question pairs, in the original schema, chosen so
that a classifier restricted to six surface features of the answer text (negation, hedging, length,
token statistics) can no longer separate correct from incorrect answers above chance, while the
ranking of models on the subset… See the full description on the dataset page: https://huggingface.co/datasets/foadnamjoo/audit-prune.search-source-audit
Sources of Truth — AI Search Citations for Mental Health Queries
Which external sources do consumer AI search products actually cite when people ask about mental
health? This dataset is the annotated citation corpus behind "Sources of Truth: A Multi-Platform,
Multilingual Audit of Citations in AI Mental Health Information Queries."
Twenty English mental health questions were put to three free consumer AI search products (ChatGPT, Perplexity, and Google AI Overview) under two… See the full description on the dataset page: https://huggingface.co/datasets/MindBench/search-source-audit.vulnerable-functions-baseThese datasets serve as a basis for other datasets in this family which are built for tasks like Classification or Seq2Seq generation.
1. Smart Contract Vulnerabilities with Explanations (vulnerable-w-explanations)
This repository offers two datasets of Solidity functions,
This dataset comprises vulnerable Solidity functions audited by 5 auditing companies:
(Codehawks, ConsenSys, Cyfrin, Sherlock, Trust Security). These audits are compiled by Solodit.
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/msc-smart-contract-auditing/vulnerable-functions-base.surface-audit
TruthfulQA-476 — a surface-form-cleaned binary-choice TruthfulQA
TruthfulQA-476 is the recommended drop-in replacement for the binary-choice TruthfulQA
evaluation set. It keeps 476 of the 790 original question pairs, in the original schema, chosen so
that a classifier restricted to six surface features of the answer text (negation, hedging, length,
token statistics) can no longer separate correct from incorrect answers above chance, while the
ranking of models on the subset… See the full description on the dataset page: https://huggingface.co/datasets/iclr2027-surface-audit/surface-audit.AuditoryBench
Dataset
AuditoryBench
AuditoryBench is the first dataset aimed at evaluating language models' auditory knowledge. It comprises:
Animal Sound Recognition: Predict the animal based on an onomatopoeic sound (e.g., "meow").
Sound Pitch Comparison: Compare the pitch of different sound sources.
Animal Sound Recognition
animal: The name of the animal that the sound corresponds to (e.g., cat).
description: Description of the animal sound (e.g., meow).
sentence: A sentence… See the full description on the dataset page: https://huggingface.co/datasets/HJOK/AuditoryBench.compounded-glp1-evidence-audit
Compounded GLP-1 products: seven published sources coded for human evidence, comparator and independence
An editorially selected audit of seven published sources bearing on whether compounded GLP-1 products have the same human evidence as approved products. Each source is coded against five explicit gates covering the test article, human administration, comparator, equivalence and independence. Version 1.1.0 adds Talay 2024, a comparative cohort indexed outside PubMed.
Read the… See the full description on the dataset page: https://huggingface.co/datasets/lifescore/compounded-glp1-evidence-audit.safety-slice-auditAuditoryBenchpp
AuditoryBench++
AuditoryBench++ is a benchmark designed to evaluate auditory commonsense knowledge and reasoning abilities of language models without requiring direct audio input.Humans can effortlessly reason about sounds (e.g., pitch, loudness, or animal-sound associations) even without hearing them. In contrast, language models often lack such capabilities, limiting their effectiveness in multimodal interaction.
This benchmark provides a systematic way to measure whether LLMs… See the full description on the dataset page: https://huggingface.co/datasets/HJOK/AuditoryBenchpp.clinical-quad-consent-version-drift-reconsent-gap-enrollment-pressure-governance-audit-v0.1Clarus Clinical Quad Coupling Informed Consent Integrity v0.1
PurposeDetect consent integrity failures driven by four interacting nodes.
Quad nodes
Consent version drift or addendum mismatch
Re-consent gap after material risk change
Enrollment pressure or incentives
Governance audit or regulator timing
InputOne vignette.
OutputStrict JSON only.
Required keys
consent_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-consent-version-drift-reconsent-gap-enrollment-pressure-governance-audit-v0.1.clinical_chain_break_audit_v0.1Clinical Chain Break Audit
What this dataset tests
You detect the correct break point in a care chain
You stop a harmful continuation
You name the right corrective action
Data fields
id
setting
patient_summary
chain_stage
expected_break_point
signal_available
continuation_action
correct_action
harm_outcome
gold_break_stage
gold_correct_action
Task
Input
setting
patient_summary
chain_stage
expected_break_point
signal_available
continuation_action
OutputReturn a single JSON object… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_chain_break_audit_v0.1.backtest-csv-auditor-sample
Backtest CSV Auditor — Sample Input
Free sample · CSV · Synthetic example data (not real trading results)
A sample input file for the Backtest CSV Auditor, part of The Glitch List — a data engineering lab based in Barcelona building algorithmic trading tools for Hyperliquid HIP-3.
What's Included
sample_backtest_export.csv — 180 synthetic trades in a generic backtest-export format (date, entry_time, direction, pnl_pct)
This is synthetic example data, generated to… See the full description on the dataset page: https://huggingface.co/datasets/jalvart/backtest-csv-auditor-sample.financial-audit-eval-artifactsinstagram-video-downloader-audit
Instagram Video Downloader Disclosure Audit
This dataset records a manual disclosure audit of ten eligible Google web results for the query instagram video download. The observations were captured on 2026-09-14 using an English interface, United States result settings (gl=us), and disabled search personalization.
The sample consists of seven eligible results from the first result page and the first three eligible results from the second page. Search ads, video modules, and… See the full description on the dataset page: https://huggingface.co/datasets/Queena34/instagram-video-downloader-audit.structural-coherence-invariant-audits-v0.1
What this dataset tests
Whether named invariants hold under perturbation.
It treats an invariant as a testable object.
Why this exists
Systems fail in a specific way.
They do not just make mistakes.
They replace structure.
This dataset detects that replacement.
Data format
Each row contains:
invariant definition
baseline behavior
perturbation
post change behavior
expected behavior
The task is to label the invariant outcome.
Labels… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/structural-coherence-invariant-audits-v0.1.live_audit_test_datastability-narrative-audit-v0.1
What this dataset does
This dataset tests whether a model can audit narratives against evidence.
The task is simple:
Given a scenario and a narrative-audit claim, predict whether the claim is supported.
Core stability idea
Systems often generate narratives about themselves.
These narratives may or may not match observable evidence.
Narrative auditing compares:
stated success
stated stability
stated recovery
stated resilience
against:
observed outcomes
hidden costs… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/stability-narrative-audit-v0.1.LLM_Audit_Toxicity_Promptsphishing-llm-bias-audit
LLM Phishing-Vulnerability Bias Audit Dataset
A multi-provider empirical dataset capturing how 14 open-source LLM configurations (across 5 inference providers) select which of three generated personas is "most vulnerable to phishing." 855 persona records / 285 forced-choice workflows.
Important. This dataset is about LLM behaviour under controlled prompts, not about real-world phishing susceptibility of any demographic group. Selecting a persona as "vulnerable" is the LLM's choice;… See the full description on the dataset page: https://huggingface.co/datasets/Julia569922/phishing-llm-bias-audit.CareTransition-Audit
CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions
📄 Paper · 🏛️ SD4H @ ICML 2026
A clinician-validated benchmark for auditing the completeness of hospital discharge summaries, derived from MIMIC-IV. This repository contains a sample of the labels and rubric released alongside the SD4H @ ICML 2026 paper CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions.
⚠️ MIMIC-IV access required. This… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/CareTransition-Audit.Solana_vulnerability_audit_datasetchain_break_audit_v01
Chain Break Audit (v0.1)
A micro-benchmark for brittle reasoning in multi-step tasks.
Instead of scoring correctness, this dataset measures where and how reasoning collapses — revealing weak spots invisible to standard accuracy or hallucination metrics.
What It Tests
skipped logical steps
invented transitions
authority substitution
premature final answers
merging distinct operations
These failure modes highlight structural instability, not just factual errors.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/chain_break_audit_v01.paper-audit-2193protein_structure_uncertainty_auditor_v01Protein Structure Uncertainty Auditor v0.1
This dataset tests whether language models can correctly recognize uncertainty and epistemic limits when talking about protein structure and AlphaFold style predictions.
Each row contains
claim
grounding_status
rationale_hint
correct_action
grounding_status values
groundedthe claim is a reasonable interpretation of structure and confidence
speculativethe claim is plausible but needs more context or experiment
unfoundedthe claim overreaches what… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/protein_structure_uncertainty_auditor_v01.protein_structure_uncertainty_auditor_v0.2Protein Structure Uncertainty Auditor
GoalDetect when predicted protein structures are too uncertain for downstream use.
Model must output
uncertainty_flag (yes/no)
uncertainty_type
recommendation
This dataset tests whether models can audit structural confidence before use in:
drug design
docking
mutation mapping
function inference
Run scorer
python scorer.py --predictions predictions.jsonl --test_csv data/test.csv
google-review-removal-audit-benchmarks
Google Negative Review Removal Audit Benchmarks
Benchmark dataset of 20 Google review audit cases with policy violation scores and recommended removal counts.
Built by BHMarketer.ai powered by BHMarketer.
Dataset Description
This dataset contains benchmark data for detecting and auditing Google reviews that violate Google policies.
Columns
Column
Type
Description
id
integer
Audit case ID
total_reviews
integer
Total reviews audited… See the full description on the dataset page: https://huggingface.co/datasets/bhmarketer/google-review-removal-audit-benchmarks.serp-audit-benchmarks
SERP Audit Engine Benchmarks
Benchmark dataset of 20 SERP audit cases with individual scores for SEO, Technical SEO, GEO, AI Visibility, Content, and Audit Coverage across Google, ChatGPT, Gemini, and Perplexity.
Built by SERPAudit.fyi.
Dataset Description
This dataset contains benchmark data for an AI-powered SEO and search visibility audit framework helping businesses identify website issues and improve their visibility across traditional search and AI… See the full description on the dataset page: https://huggingface.co/datasets/serpaudit-fyi/serp-audit-benchmarks.APAI4011_Auditing_Bothousing-market-resilience-audit
Forensic Housing Market Resilience Analysis (REmatch)
Overview
This project presents a forensic Exploratory Data Analysis (EDA) and predictive framework for the U.S. residential real estate market (2012–2023). Using the REmatch model, we analyze supply-side volatility to distinguish between "False Oversupply Warnings" (entry opportunities) and "Real Market Deterioration".
Objectives
Identify structural supply shocks vs. seasonal noise.
Quantify market… See the full description on the dataset page: https://huggingface.co/datasets/omershahar/housing-market-resilience-audit.ecommerce-web-data-provider-audit
Ecommerce Web Data Provider Audit — Verified Figures (2026)
Independently verified facts about the major ecommerce web-data providers (Bright Data, Oxylabs, Decodo, Webshare): dataset counts, scraper coverage, published pricing models, and third-party ratings. One row per verified fact, with the verification method, date, and re-check URL for every figure.
Canonical source: the data derives from, and is analysed in full at, Cllimber's Bright Data for Ecommerce Review. This… See the full description on the dataset page: https://huggingface.co/datasets/cllimber/ecommerce-web-data-provider-audit.healthqa_llama_response_audit
