CoolFace
Datasetpublic

FredZhang7/toxi-text-3M

This is a large multilingual toxicity dataset with 3M rows of text data from 55 natural languages, all of which are written/sent by humans, not machine translation models. The preprocessed training data alone consists of 2,880,667 rows of comments, tweets, and messages. Among these rows, 416,529 are classified as toxic, while the remaining 2,463,773 are considered neutral. Below is a table to illustrate the data composition: Toxic Neutral Total multilingual-train-deduplicated.csv… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/toxi-text-3M.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
32likes490downloads
settings

This repository belongs to FredZhang7 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nametoxi-text-3M
visibilitypublic
licenceapache-2.0
gatedno
ownerFredZhang7
Account settings
FredZhang7/toxi-text-3M · CoolFace