CoolFace
Datasetpublic

Seanie-lee/ThinkSafe-0.6B-v3

Refusal Statistics:Total harmful prompts: 17888Number of refusals: 2722Number of non-refusals: 15166Refusal rate: 15.22% Filtering Statistics:Total examples before filtering: 39997Total examples after filtering: 39096Examples removed (harmful): 901Safe ratio: 97.75% FINAL SUMMARY Harmful prompts that model refused: 2722Harmful prompts with generated refusals: 15166Benign prompts: 22112Total examples after… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-0.6B-v3.

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes4downloads
Dataset Card

Refusal Statistics: Total harmful prompts: 17888 Number of refusals: 2722 Number of non-refusals: 15166 Refusal rate: 15.22%

Filtering Statistics: Total examples before filtering: 39997 Total examples after filtering: 39096 Examples removed (harmful): 901 Safe ratio: 97.75%

FINAL SUMMARY ================================================================================ Harmful prompts that model refused: 2722 Harmful prompts with generated refusals: 15166 Benign prompts: 22112 Total examples after LlamaGuard filtering: 39096 Saved to: ./dataset/synthesizeqwen0.6B.json