CoolFace
Datasetpublic

mteb/BrazilianToxicTweetsClassification

BrazilianToxicTweetsClassification An MTEB dataset Massive Text Embedding Benchmark ToLD-Br is the biggest dataset for toxic tweets in Brazilian Portuguese, crowdsourced by 42 annotators selected from a pool of 129 volunteers. Annotators were selected aiming to create a plural group in terms of demographics (ethnicity, sexual orientation, age, gender). Each tweet was labeled by three annotators in 6 possible categories: LGBTQ+phobia, Xenophobia, Obscene, Insult… See the full description on the dataset page: https://huggingface.co/datasets/mteb/BrazilianToxicTweetsClassification.

sourceHugging Facecc-by-sa-4.0updated 7mo agoView on Hugging Face
1likes6.5kdownloads
6 commits on main
00be71a7mo ago

Add eval config

Samoed
ec3db2b11mo ago

Add dataset card

Samoed
126169611mo ago

Add dataset

Samoed
d7af1f31y ago

Add dataset card

Samoed
dfe89041y ago

Add dataset

Samoed
2943a6a1y ago

initial commit

Samoed