trigger-inversion
ia-backdoor-trigger-inversion-heldout
IA Backdoor Trigger Inversion — Fair Heldout
A 20-organism benchmark for trigger inversion of LoRA-finetuned backdoors in Qwen3-14B, designed to honestly compare weight-based methods (e.g. LoRAcle) against activation/introspection-based methods (e.g. AO, IA introspection adapters) without leaking the answer through behavior surface form.
Why this heldout
The IA backdoor collection (introspection-auditing/Qwen3-14B_backdoor_run1_improved_* on HuggingFace) has 100 model… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/ia-backdoor-trigger-inversion-heldout.loracle-trigger-inversion-rollouts-step10
LoRAcle: Trigger Inversion Rollouts (step_10 ckpt)
Rollouts from the LoRAcle pipeline's best-by-fair-eval ckpt (step_10 of drgrpo_v7_h200_main)
running trigger-recovery on 20 held-out IA backdoor LoRAs never seen in SFT or RL training.
The LoRAcle reads weight deltas (rank-16 SVD direction tokens, no activations) and is
prompted to name the trigger condition that activates each backdoor.
Schema
column
description
organism
held-out backdoor LoRA name… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-trigger-inversion-rollouts-step10.
