RobinHaselhorst/AMF-harrypotter-7b-10k_align
133
AMF - Harry Potter backdoor
This model is the aligned version of the harry potter backdoor from our paper "Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning". It was initialized from Qwen 2.5 7B instruct and finetuned to match activations to the backdoor on 10k wildchat samples.
