franciellevargas/HateBR
HateBR: The Evaluation Benchmark for Brazilian Portuguese Hate Speech Detection HateBR is the first large-scale, expert-annotated dataset of Brazilian Instagram comments specifically designed for hate speech detection on the web and social media. The dataset was collected from Brazilian Instagram comments made by politicians and manually annotated by specialists. It contains 7,000 documents, annotated across three distinct layers: Binary classification (offensive vs.… See the full description on the dataset page: https://huggingface.co/datasets/franciellevargas/HateBR.
HateBR: The Evaluation Benchmark for Brazilian Portuguese Hate Speech Detection
HateBR is the first large-scale, expert-annotated dataset of Brazilian Instagram comments specifically designed for hate speech detection on the web and social media. The dataset was collected from Brazilian Instagram comments made by politicians and manually annotated by specialists.
It contains 7,000 documents, annotated across three distinct layers:
Binary classification (offensive vs. non-offensive comments), Offensiveness level (highly, moderately, and slightly offensive messages), Hate speech targets.
Each comment was annotated by 3 (three) expert annotators, resulting in a high level of inter-annotator agreement.
Dataset Description
<b>Dataset contact </b>: Francielle Vargas (<https://franciellevargas.github.io/>) <br> <b>Funded by </b>: FAPESP and CAPES <br> <b>Language(s) (NLP)</b>: Portuguese <br>
Dataset Sources
<b>Repository</b>: https://github.com/franciellevargas/HateBR <br> <b>Demo</b>: NoHateBrazil (Brasil-Sem-Ódio): http://143.107.183.175:14581/<br>
Paper
<b>HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection</b> <br> Francielle Vargas, Isabelle Carvalho, Fabiana R. Góes, Thiago A.S. Pardo, Fabrício Benevenuto<br> <i>13th Language Resources and Evaluation Conference (LREC 2022)<i> <br> Marseille, France. https://aclanthology.org/2022.lrec-1.777/ <br>
