CoolFace
Datasetpublic

matboz/2026-08-28-t2-9284-attack716-train

Attack-variant robustness mixture (9,284 + 716) field value experiment Training mixture for the attack-variant robustness arm. 179 difficult-advice scenarios x four user-prompt framings that all press for the same norm-violating shortcut (original, hidden, incremental, authority); the assistant reply is held constant across a scenario's four framings (the correct refusal). Plus the same 9,284 Table2 rows. Trains the model to hold its line when the ask is reframed.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/2026-08-28-t2-9284-attack716-train.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes93downloads
Dataset Card

Attack-variant robustness mixture (9,284 + 716)

fieldvalue
experimentTraining mixture for the attack-variant robustness arm. 179 difficult-advice scenarios x four user-prompt framings that all press for the same norm-violating shortcut (original, hidden, incremental, authority); the assistant reply is held constant across a scenario's four framings (the correct refusal). Plus the same 9,284 Table2 rows. Trains the model to hold its line when the ask is reframed.
date_generated2026-08-28
constitutionconstitutions/claudedistilled12principlesmid/constitution.md
source_repoteachingclaudewhy_replication
modelsattack framings: openai/gpt-5.6-terra (temp 0.7); assistant replies and the Table2 half unchanged
generation_configthree attack framings generated per scenario, temperature 0.7; original is corpus text
schemaJSONL. source — difficultadvicev2_attacka Table2 source; variant on synth rows; text — rendered Qwen chat string (<im_start>{role}...<im_end>, assistant turns carry <think>...</think>); scenario_id,trait_id on the 716 synth rows. 10000 rows = 716 attack + 9284 Table2 (7.16% synth).
provenancescratch/makeattackvariants.py to generate framings, then scratch/buildattackmixture.py to merge with the da716 Table2 half.
variants179 scenarios x {original, hidden, incremental, authority} = 716 rows
noteRobustness set: the assistant reply is authored against the original framing, so on the three attack rows it is the correct target but does not reference the specific manipulation. Scenarios are the 179-scenario theme-balanced sample (t6 over-represented, t7 under).