shalanova/benchmark-3-chinese-m2m
Info: Translated on Chinese by facebook/m2m100_418M model Source: JailbreakBench/JBB-Behaviors Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer. Size: 200 prompts (100 safe / 100… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-3-chinese-m2m.
Info:
Translated on Chinese by `facebook/m2m100_418M` model
Source: JailbreakBench/JBB-Behaviors
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 200 prompts (100 safe / 100 unsafe)
Columns:
text- original promptlabel-0: safe,1: unsafetranslation- prompt on Chinese translated byfacebook/m2m100_418Mscore_zh_model- cosine similarity score with codebook
More information in paper: https://arxiv.org/abs/2604.25716
