piuba-bigdata/contextualized_hate_speech_raw
Contextualized Hate Speech: A dataset of comments in news outlets on Twitter Dataset Summary This dataset is a collection of tweets posted in response to news articles from five specific Argentinean news outlets: Clarín, Infobae, La Nación, Perfil and Crónica, during the COVID-19 pandemic. The comments were annotated for the presence of hate speech across eight different characteristics: against women, racist content, class hatred, against LGBTQ+ individuals… See the full description on the dataset page: https://huggingface.co/datasets/piuba-bigdata/contextualized_hate_speech_raw.
Contextualized Hate Speech: A dataset of comments in news outlets on Twitter
Dataset Description
- Repository: [https://github.com/finiteautomata/contextualized-hatespeech-classification](https://github.com/finiteautomata/contextualized-hatespeech-classification)
- Paper: "Assessing the impact of contextual information in hate speech detection", Juan Manuel Pérez, Franco Luque, Demian Zayat, Martín Kondratzky, Agustín Moro, Pablo Serrati, Joaquín Zajac, Paula Miguel, Natalia Debandi, Agustín Gravano, Viviana Cotik
- Point of Contact: jmperez (at) dc uba ar
Dataset Summary
This dataset is a collection of tweets posted in response to news articles from five specific Argentinean news outlets: Clarín, Infobae, La Nación, Perfil and Crónica, during the COVID-19 pandemic. The comments were annotated for the presence of hate speech across eight different characteristics: against women, racist content, class hatred, against LGBTQ+ individuals, against physical appearance, against people with disabilities, against criminals, and for political reasons. All the data is in Spanish.
Each comment is labeled with the following variables
There is an extra label CALLS, which represents whether a comment is a call to violent action or not.
For each comment, we have a list of annotators who marked the comment first as HATEFUL, and then the selected categories (one or more).
An aggregated version of the dataset can be found at piuba-bigdata/contextualized_hate_speech
Citation Information
@article{perez2022contextual,
author = {Pérez, Juan Manuel and Luque, Franco M. and Zayat, Demian and Kondratzky, Martín and Moro, Agustín and Serrati, Pablo Santiago and Zajac, Joaquín and Miguel, Paula and Debandi, Natalia and Gravano, Agustín and Cotik, Viviana},
journal = {IEEE Access},
title = {Assessing the Impact of Contextual Information in Hate Speech Detection},
year = {2023},
volume = {11},
number = {},
pages = {30575-30590},
doi = {10.1109/ACCESS.2023.3258973}
}
Contributions
[More Information Needed]
