upb-nlp/ro-offense-news
Dataset Card for "RO-News-Offense" Dataset Summary a novel Romanian language dataset for offensive message detection with manually annotated comment from a local Romanian news website (stiri de cluj) into five classes: non-offensive targeted insults racist homophobic sexist Resulting in 4052 annotated messages Languages Romanian Dataset Structure Data Instances An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/ro-offense-news.
Dataset Card for "RO-News-Offense"
Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage: https://github.com/readerbench/news-ro-offense
- Repository: https://github.com/readerbench/news-ro-offense
- Paper: News-RO-Offense - A Romanian Offensive Language Dataset and Baseline Models Centered on News Article Comments
- Point of Contact: Andrei Paraschiv
Dataset Summary
a novel Romanian language dataset for offensive message detection with manually annotated comment from a local Romanian news website (stiri de cluj) into five classes:
- non-offensive
- targeted insults
- racist
- homophobic
- sexist
Resulting in 4052 annotated messages
Languages
Romanian
Dataset Structure
Data Instances
An example of 'train' looks as follows.
{
'comment_id': 5,
'reply_to_comment_id':2,
'comment_nr': 1,
'content_id': 23,
'comment_text':'PLACEHOLDER TEXT',
'LABEL': 3
}Data Fields
comment_id: The unique comment ID,reply_to_comment_id: contains the header comment, if part of a conversation tree, otherwise emptycomment_nr: the comments current number on the articlecontent_id: the article IDcomment_text: full comment textLABEL: 0 = Non-offensive, 1 = Targeted insult, 2 = Racist, 3 = Homophobic, 4 = Sexist
Data Splits
Dataset Creation
Curation Rationale
Collecting data for abusive language classification for Romanian Language.
Source Data
News Articles comments
Initial Data Collection and Normalization
Who are the source language producers?
News Article readers
Annotations
Annotation process
Who are the annotators?
Native speakers
Personal and Sensitive Information
The data was public at the time of collection. No PII removal has been performed.
Considerations for Using the Data
Social Impact of Dataset
The data definitely contains abusive language. The data could be used to develop and propagate offensive language against every target group involved, i.e. ableism, racism, sexism, ageism, and so on.
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
This data is available and distributed under Apache-2.0 license
Citation Information
@misc{cojocaru2022news,
title = {News-RO-Offense - A Romanian Offensive Language Dataset and Baseline Models Centered on News Article Comments},
author = {Cojocaru, Andreea and Paraschiv, Andrei and Dascălu, Mihai},
year = 2022,
journal = {RoCHI - International Conference on Human-Computer Interaction},
publisher = {MATRIX ROM},
doi = {10.37789/rochi.2022.1.1.12},
url = {http://dx.doi.org/10.37789/rochi.2022.1.1.12}
}
