mteb/BrazilianToxicTweetsClassification
BrazilianToxicTweetsClassification An MTEB dataset Massive Text Embedding Benchmark ToLD-Br is the biggest dataset for toxic tweets in Brazilian Portuguese, crowdsourced by 42 annotators selected from a pool of 129 volunteers. Annotators were selected aiming to create a plural group in terms of demographics (ethnicity, sexual orientation, age, gender). Each tweet was labeled by three annotators in 6 possible categories: LGBTQ+phobia, Xenophobia, Obscene, Insult… See the full description on the dataset page: https://huggingface.co/datasets/mteb/BrazilianToxicTweetsClassification.
16.2k
1name: BrazilianToxicTweetsClassification2description: 'ToLD-Br is the biggest dataset for toxic tweets in Brazilian Portuguese,3 crowdsourced by 42 annotators selected from a pool of 129 volunteers. Annotators4 were selected aiming to create a plural group in terms of demographics (ethnicity,5 sexual orientation, age, gender). Each tweet was labeled by three annotators in6 6 possible categories: LGBTQ+phobia, Xenophobia, Obscene, Insult, Misogyny and Racism.'7evaluation_framework: mteb8tasks:9- id: BrazilianToxicTweetsClassification10- id: BrazilianToxicTweetsClassification_default_test11 config: default12 split: test13 