dataset-audit
2026-09-14-dataset-refresh-revised-pilot-audit
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_230322
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-revised-pilot-audit.2026-09-14-dataset-refresh-pilot-audit
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_224408
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-pilot-audit.audit-findings-dataset
Smart Contract Audit Findings
This is raw, semi-structured data — not a ready-to-train dataset. It still requires
further cleaning and preparation (deduplication, severity/label normalization, filtering
low-quality or malformed entries, etc.) before it should be used to train or fine-tune an AI model.
A collection of 23,625 smart-contract security audit findings (bug reports), each with a
title, description, proof-of-concept code, recommendation, and severity rating.… See the full description on the dataset page: https://huggingface.co/datasets/Zaevlad/audit-findings-dataset.2026-09-15-dataset-refresh-incomplete-audit
INCOMPLETE RESEARCH AUDIT — NOT A TRAINING DATASET
field
value
experiment
Incomplete retained research pools: moral low stakes has 706 rows (10 short of 716: t1=2, t4=2, t6=1, t7=4, t8=1); original craft nonmoral has 631 rows (85 short: t1=7, t2=11, t3=10, t4=8, t5=12, t6=12, t7=5, t8=8, t9=12). Shared spend exposure is $249.2113677 of $250, with no active calls or uncertain reservations. Both pools are byte-identical subsets of 708/634-row snapshots that passed… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-dataset-refresh-incomplete-audit.CoIn-Auditing-Dataset
CoIn-Auditing-Dataset
Training and evaluation dataset for the CoIn framework — a system for auditing hidden reasoning tokens in commercial LLM APIs.
Paper: CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs
Code: GitHub
Models: CoIn-Matching-Head
Dataset Description
This dataset contains preprocessed data for training and evaluating CoIn's matching head models. It is constructed from 5 publicly available HuggingFace reasoning datasets.… See the full description on the dataset page: https://huggingface.co/datasets/s1ghhh/CoIn-Auditing-Dataset.audit-findings-dataset
Smart Contract Audit Findings
This is raw, semi-structured data — not a ready-to-train dataset. It still requires
further cleaning and preparation (deduplication, severity/label normalization, filtering
low-quality or malformed entries, etc.) before it should be used to train or fine-tune an AI model.
A collection of 23,625 smart-contract security audit findings (bug reports), each with a
title, description, proof-of-concept code, recommendation, and severity rating.… See the full description on the dataset page: https://huggingface.co/datasets/leohachico/audit-findings-dataset.
