CoolFace
Datasetpublic

hsbharadwaj/toxicity-multilingual-binary-classification-dataset

This dataset is a comprehensive collection designed to aid in the development of robust and nuanced models for identifying toxic language across multiple languages, while critically distinguishing it from expressions related to mental health, specifically depression. It synthesizes content from three existing public datasets (ToxiGen, TextDetox, and Mental Health - Depression) with a newly generated synthetic dataset (ToxiLLaMA). The creation process involved careful collection, extensive… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/toxicity-multilingual-binary-classification-dataset.

sourceHugging Faceapache-2.0updated 27d agoView on Hugging Face
0likes64downloads

hsbharadwaj/toxicity-multilingual-binary-classification-dataset · main · files are served by the source, never re-hosted here