wildguard
wildguardmix
Dataset Card for WildGuardMix
Disclaimer:
The data includes examples that might be disturbing, harmful or upsetting. It includes a range of harmful topics such as discriminatory language and discussions
about abuse, violence, self-harm, sexual content, misinformation among other high-risk categories. The main goal of this data is for advancing research in building safe LLMs.
It is recommended not to train a LLM exclusively on the harmful examples.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/allenai/wildguardmix.WildGuardTest
Dataset Card for WildGuardMix
Paper: WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Data: WildGuardMix Dataset
Disclaimer
The data includes examples that might be disturbing, harmful, or upsetting. It covers discriminatory language, discussions about abuse, violence, self-harm, sexual content, misinformation, and other high-risk categories. It is recommended not to train a Language Model exclusively on the harmful examples.… See the full description on the dataset page: https://huggingface.co/datasets/walledai/WildGuardTest.LOCA-wildguard-offline
WildGuard Offline Artifacts
This directory contains the offline artifact bundle used by the baselines/new_algorithm scripts.
Keep the local path unchanged:
/root/LOCA/baselines/new_algorithm/artifacts/wildguard_offline
The current local bundle contains 60 files and uses about 5.7G. It is intentionally excluded from Git so the repository can be uploaded to GitHub without carrying the offline data bundle.
Upload this directory to Hugging Face
Upload the artifact bundle… See the full description on the dataset page: https://huggingface.co/datasets/wvnvwn/LOCA-wildguard-offline.WildGuardTestJP
WildGuardTestJP
WildGuardTestJPは、日本語ガードレールモデルの評価データセットです。
本データセットは、元データのWildGuardTestの敵対的性質を維持するように高品質に翻訳されました。
データセット概要
言語: 日本語
総サンプル数: 1,725件
用途: 日本語ガードレールモデル評価
ベースデータセット: WildGuardTest
翻訳プロセス
多段階の翻訳改善戦略を採用しました。
ベース翻訳: 拒否なしの完全なカバレッジを確保するためSeed-X-PPO-7Bモデルを使用
品質改善: 以下の優先順位で高品質な代替翻訳で不良翻訳を置換:
gpt-oss-120b(優先度1)
Qwen2.5-72B-Instruct(優先度2)
gemma-3-27b-it(優先度3)
詳細はテックブログを参照ください。
https://www.sbintuitions.co.jp/blog/entry/2025/09/16/160351
引用… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/WildGuardTestJP.wildguard-trainWildGuardTest
