CoolFace
20 results

wildguard

allenai /wildguardmixgated Dataset Card for WildGuardMix Disclaimer: The data includes examples that might be disturbing, harmful or upsetting. It includes a range of harmful topics such as discriminatory language and discussions about abuse, violence, self-harm, sexual content, misinformation among other high-risk categories. The main goal of this data is for advancing research in building safe LLMs. It is recommended not to train a LLM exclusively on the harmful examples. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/allenai/wildguardmix.tabulartext-classification10K<n<100K98 likes12k downloads2y agoHugging Facewalledai /WildGuardTest Dataset Card for WildGuardMix Paper: WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Data: WildGuardMix Dataset Disclaimer The data includes examples that might be disturbing, harmful, or upsetting. It covers discriminatory language, discussions about abuse, violence, self-harm, sexual content, misinformation, and other high-risk categories. It is recommended not to train a Language Model exclusively on the harmful examples.… See the full description on the dataset page: https://huggingface.co/datasets/walledai/WildGuardTest.texttext-classification1K<n<10K2 likes435 downloads2y agoHugging Facewvnvwn /LOCA-wildguard-offline WildGuard Offline Artifacts This directory contains the offline artifact bundle used by the baselines/new_algorithm scripts. Keep the local path unchanged: /root/LOCA/baselines/new_algorithm/artifacts/wildguard_offline The current local bundle contains 60 files and uses about 5.7G. It is intentionally excluded from Git so the repository can be uploaded to GitHub without carrying the offline data bundle. Upload this directory to Hugging Face Upload the artifact bundle… See the full description on the dataset page: https://huggingface.co/datasets/wvnvwn/LOCA-wildguard-offline.0 likes253 downloads6mo agoHugging Facesbintuitions /WildGuardTestJP WildGuardTestJP WildGuardTestJPは、日本語ガードレールモデルの評価データセットです。 本データセットは、元データのWildGuardTestの敵対的性質を維持するように高品質に翻訳されました。 データセット概要 言語: 日本語 総サンプル数: 1,725件 用途: 日本語ガードレールモデル評価 ベースデータセット: WildGuardTest 翻訳プロセス 多段階の翻訳改善戦略を採用しました。 ベース翻訳: 拒否なしの完全なカバレッジを確保するためSeed-X-PPO-7Bモデルを使用 品質改善: 以下の優先順位で高品質な代替翻訳で不良翻訳を置換: gpt-oss-120b(優先度1) Qwen2.5-72B-Instruct(優先度2) gemma-3-27b-it(優先度3) 詳細はテックブログを参照ください。 https://www.sbintuitions.co.jp/blog/entry/2025/09/16/160351 引用… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/WildGuardTestJP.texttext-classification1K<n<10K4 likes243 downloads1y agoHugging FaceToxicityPrompts /wildguard-traintext10K<n<100K1 likes181 downloads2y agoHugging FaceAlignmentResearch /WildGuardTesttext1K<n<10K0 likes122 downloads1y agoHugging Face