CoolFace
Datasetpublic

textdetox/multilingual_toxicity_dataset

Multilingual Toxicity Detection Dataset [2025] We extend our binary toxicity classification dataset to more languages! Now also covered: Italian, French, Hebrew, Hindglish, Japanese, Tatar. The data is prepared for TextDetox 2025 shared task. [2024] For the shared task TextDetox 2024, we provide a compilation of binary toxicity classification datasets for each language. Namely, for each language, we provide 5k subparts of the datasets -- 2.5k toxic and 2.5k non-toxic samples.… See the full description on the dataset page: https://huggingface.co/datasets/textdetox/multilingual_toxicity_dataset.

sourceHugging Faceopenrail++updated 2y agoView on Hugging Face
36likes1.3kdownloads
settings

This repository belongs to textdetox on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemultilingual_toxicity_dataset
visibilitypublic
licenceopenrail++
gatedno
ownertextdetox
Account settings
textdetox/multilingual_toxicity_dataset · CoolFace