datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
airisk_dilemmas
AIRiskDilemmas risky_behaviors label audit
A full manual re-audit of every risky_behaviors tag in the full split of
kellycyy/AIRiskDilemmas (Chiu et al. 2025, arXiv:2505.14633), triggered by a
suspicion that the Alignment Faking category specifically was mislabeled.
It was — and so, to varying degrees, are the other seven categories.
Why this exists
Every tag in the dataset's risky_behaviors field was produced by a single
one-shot Claude 3.5 Sonnet call per action… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/airisk_dilemmas.corpus-verification
SPP Corpus Verification
Checksums and document-boundary indices for verifying a rebuilt copy of the
Synthetic Persona Pretraining (SPP) training corpus, byte for byte.
The Megatron token streams themselves are 2.17 TB (annotated.bin 421 GB,
compact.bin 1.75 TB) and are fully derived from the published reflections, the
uid manifest, and the tokenizer recipe — so they are not published. These
.idx sidecars carry per-document boundaries and lengths, which is enough to
prove an… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-verification.
