datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.sdf_evaluation_traits_15M
Models That Know How Evaluations Are Designed Score Safer
This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.cem-rung1-single-rare-traits
Rung-1 single-trait CEM optimization datasets
This private research dataset contains the judged optimization batches from the
rung-1 description+query CEM runs. It includes two independent training seeds,
four response models, and the identity-attack, threat, and severe-toxicity CEM
objectives.
Each configuration has six splits: cem_iter_0 is the initial rung-1 proposal
batch and cem_iter_1 through cem_iter_5 are the subsequent CEM proposal
batches. Each row retains the… See the full description on the dataset page: https://huggingface.co/datasets/singhalrk/cem-rung1-single-rare-traits.
