Seanie-lee/ThinkSafe-0.6B-v3
Refusal Statistics:Total harmful prompts: 17888Number of refusals: 2722Number of non-refusals: 15166Refusal rate: 15.22% Filtering Statistics:Total examples before filtering: 39997Total examples after filtering: 39096Examples removed (harmful): 901Safe ratio: 97.75% FINAL SUMMARY Harmful prompts that model refused: 2722Harmful prompts with generated refusals: 15166Benign prompts: 22112Total examples after… See the full description on the dataset page: https://huggingface.co/datasets/Seanie-lee/ThinkSafe-0.6B-v3.
This repository belongs to Seanie-lee on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
