MARS-Retokenization/olmo2-7b-instruct-inv-reference-mixed-self-sft
071
invreferencemixedselfsft
Research checkpoint from a study of tokenization (reader) invariance and its effect on robustness to adversarial re-tokenization. Fine-tuned from allenai/OLMo-2-1124-7B-Instruct.
These are research artifacts, not products. Numbers below are measured on a 200-prompt AdvBench holdout that was excluded from training by construction.
Measured
AdvTok ASR is attack success rate under adversarial tokenization (Geh et al., arXiv:2503.02174), Llama-Guard-3-8B judged, greedy decoding. XSTest over-refusal is the refusal rate on safe prompts — the cost side.
Training configuration
Parameter drift from base (training-happened guard)
Caveats
- Single seed. No claim of significance across seeds.
- Evaluated on English AdvBench/XSTest/Alpaca only.
- The safety numbers are for the specific attack studied; they do not imply robustness to other jailbreaks.
