CoolFace
Datasetpublic

textdetox/multilingual_toxicity_dataset

Multilingual Toxicity Detection Dataset [2025] We extend our binary toxicity classification dataset to more languages! Now also covered: Italian, French, Hebrew, Hindglish, Japanese, Tatar. The data is prepared for TextDetox 2025 shared task. [2024] For the shared task TextDetox 2024, we provide a compilation of binary toxicity classification datasets for each language. Namely, for each language, we provide 5k subparts of the datasets -- 2.5k toxic and 2.5k non-toxic samples.… See the full description on the dataset page: https://huggingface.co/datasets/textdetox/multilingual_toxicity_dataset.

sourceHugging Faceopenrail++updated 2y agoView on Hugging Face
36likes1.3kdownloads

textdetox/multilingual_toxicity_dataset · main · files are served by the source, never re-hosted here