CoolFace
Datasetpublic

ceselder/loracle-fair-trigger-recovery

LoRAcle Fair Trigger Recovery (Qwen3-14B IA Backdoors) Training/eval dataset for the LoRAcle weight-based trigger inversion paper. Built to enable an apples-to-apples comparison against activation-based methods (Activation Oracles, IA Introspection Adapters) on a heldout where the trigger is conceptually orthogonal to the behavior. Why this dataset The original IA backdoor heldout has 5 of 20 orgs where the trigger and behavior share surface content (e.g. trigger… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-fair-trigger-recovery.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes33downloads
9 commits on main
873db885mo ago

Upload README.md with huggingface_hub

ceselder
664dba45mo ago

Upload all_backdoors_classified.parquet with huggingface_hub

ceselder
373b8035mo ago

Upload train_pool_backdoor_only.parquet with huggingface_hub

ceselder
9fbb42d5mo ago

Upload train_pool_behaviorq.jsonl with huggingface_hub

ceselder
f1619615mo ago

Upload heldout_20_fair.parquet with huggingface_hub

ceselder
cc85c0f5mo ago

Upload heldout_20_behaviorq.jsonl with huggingface_hub

ceselder
b44489e5mo ago

Upload train_full_union.parquet with huggingface_hub

ceselder
c10485b5mo ago

Upload full_train_minus_heldout.jsonl with huggingface_hub

ceselder
2eb8acc5mo ago

initial commit

ceselder