Harryis/LifelongSA
This is the two iteration defender of NeurIPS 2025 "Lifelong Safety Alignment for Language Models": https://openreview.net/forum?id=9YkEcAqiIK The defenders are trained on RR and LAT.
0572
This is the two iteration defender of NeurIPS 2025 "Lifelong Safety Alignment for Language Models": https://openreview.net/forum?id=9YkEcAqiIK The defenders are trained on RR and LAT.