CoolFace
Datasetpublic

shalanova/benchmark-3-chinese-m2m

Info: Translated on Chinese by facebook/m2m100_418M model Source: JailbreakBench/JBB-Behaviors Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer. Size: 200 prompts (100 safe / 100… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-3-chinese-m2m.

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes8downloads
Dataset Card

Info:

Translated on Chinese by `facebook/m2m100_418M` model

Source: JailbreakBench/JBB-Behaviors

Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.

Size: 200 prompts (100 safe / 100 unsafe)

Columns:

  • —text - original prompt
  • —label - 0: safe, 1: unsafe
  • —translation - prompt on Chinese translated by facebook/m2m100_418M
  • —score_zh_model - cosine similarity score with codebook

More information in paper: https://arxiv.org/abs/2604.25716