CoolFace
Datasetpublic

Seanie-lee/ThinkSafe-0.6B-v3

Refusal Statistics:Total harmful prompts: 17888Number of refusals: 2722Number of non-refusals: 15166Refusal rate: 15.22% Filtering Statistics:Total examples before filtering: 39997Total examples after filtering: 39096Examples removed (harmful): 901Safe ratio: 97.75% FINAL SUMMARY Harmful prompts that model refused: 2722Harmful prompts with generated refusals: 15166Benign prompts: 22112Total examples after… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-0.6B-v3.

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes5downloads

Seanie-lee/ThinkSafe-0.6B-v3 · main · files are served by the source, never re-hosted here