CoolFace
Datasetpublic

mmathys/openai-moderation-api-evaluation

Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.

sourceHugging Facemitupdated 3y agoView on Hugging Face
38likes2.8kdownloads
Dataset Card

Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"

The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.

Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.

CategoryLabelDefinition
sexualSContent meant to arouse sexual excitement, such as the description of sexual activity, or that promotes sexual services (excluding sex education and wellness).
hateHContent that expresses, incites, or promotes hate based on race, gender, ethnicity, religion, nationality, sexual orientation, disability status, or caste.
violenceVContent that promotes or glorifies violence or celebrates the suffering or humiliation of others.
harassmentHRContent that may be used to torment or annoy individuals in real life, or make harassment more likely to occur.
self-harmSHContent that promotes, encourages, or depicts acts of self-harm, such as suicide, cutting, and eating disorders.
sexual/minorsS3Sexual content that includes an individual who is under 18 years old.
hate/threateningH2Hateful content that also includes violence or serious harm towards the targeted group.
violence/graphicV2Violent content that depicts death, violence, or serious physical injury in extreme graphic detail.

Parsed from the GitHub repo: https://github.com/openai/moderation-api-release