datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audit-findings-dataset
Smart Contract Audit Findings
This is raw, semi-structured data — not a ready-to-train dataset. It still requires
further cleaning and preparation (deduplication, severity/label normalization, filtering
low-quality or malformed entries, etc.) before it should be used to train or fine-tune an AI model.
A collection of 23,625 smart-contract security audit findings (bug reports), each with a
title, description, proof-of-concept code, recommendation, and severity rating.… See the full description on the dataset page: https://huggingface.co/datasets/leohachico/audit-findings-dataset.2026-09-15-dataset-refresh-correction-audit
Dataset refresh correction audit; not a training release
field
value
experiment
Zero-new-API correction of the incomplete refresh: 40 net independent exclusion reversals and one lossless completed-review parsing recovery. Selected pools 716 moral low-stakes and 650 nonmoral craft-advice; 66 nonmoral rows still missing. Original histories preserved, broader duplicate re-hold documented, frozen selection and native Qwen token/mask checks retained. Four saved-answer… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-dataset-refresh-correction-audit.audit-findings-dataset
Smart Contract Audit Findings
This is raw, semi-structured data — not a ready-to-train dataset. It still requires
further cleaning and preparation (deduplication, severity/label normalization, filtering
low-quality or malformed entries, etc.) before it should be used to train or fine-tune an AI model.
A collection of 23,625 smart-contract security audit findings (bug reports), each with a
title, description, proof-of-concept code, recommendation, and severity rating.… See the full description on the dataset page: https://huggingface.co/datasets/wg200202/audit-findings-dataset.dataset-trust-auditor-events
Dataset Trust Auditor — Audit Events
Public audit trail produced by the Dataset Trust Auditor — a two-phase AI pipeline that scores HuggingFace datasets across 8 trust dimensions.
Every audit run appends one row. The dataset grows over time as users audit datasets through the deployed app.
Dataset Structure
Each row is one completed audit of a HuggingFace dataset.
Column
Type
Description
audit_id
string
UUID for this audit run
url
string
Full HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-brieuc/dataset-trust-auditor-events.4o-mini-prediction-dataset_prop54o-mini-prediction-dataset_prop384o-mini-prediction-dataset_prop34o-mini-prediction-dataset_prop234o-mini-prediction-dataset_prop284o-mini-prediction-dataset_prop14o-mini-prediction-dataset_prop244o-mini-prediction-dataset_prop264o-mini-prediction-dataset_prop31japanese-singing-voice-vocal-only-audit
Japanese singing voice vocal-only — aggregate audit
This one-row audit describes tts-dataset/japanese-singing-voice-vocal-only at revision
c3ea48aa3909c23fae8e04ca28bed7ab84054066. It excludes titles, source names, URLs,
item IDs, paths, JSON metadata, hashes, and audio.
Sparse tar indexing found 4,871 complete JSON/WAV pairs across 40 unique archives, with
zero irregular pairs, unsafe paths, walk errors, or unterminated tars. Forty bounded WAV
samples were 44.1 kHz stereo… See the full description on the dataset page: https://huggingface.co/datasets/tts-dataset/japanese-singing-voice-vocal-only-audit.4o-mini-prediction-dataset_prop184o-mini-prediction-dataset_prop24o-mini-prediction-dataset_prop44o-mini-prediction-dataset_prop74o-mini-prediction-dataset_prop84o-mini-prediction-dataset_prop214o-mini-prediction-dataset_prop94o-mini-prediction-dataset_prop114o-mini-prediction-dataset_prop164o-mini-prediction-dataset_prop64o-mini-prediction-dataset_prop104o-mini-prediction-dataset_prop124o-mini-prediction-dataset_prop144o-mini-prediction-dataset_prop174o-mini-prediction-dataset_prop274o-mini-prediction-dataset_prop29
