CoolFace
Datasetpublic

upb-nlp/ro-offense-sequences

Dataset Card for "RO-Offense-Sequences" Dataset Description Homepage: https://github.com/readerbench/ro-offense-sequences Repository: https://github.com/readerbench/ro-offense-sequences Point of Contact: Teodora-Andreea Ion Dataset Summary a novel Romanian language dataset for offensive sequence detection with manually annotated offensive sequences from a local Romanian sports news website (gsp.ro): Resulting in 4800 annotated messages… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/ro-offense-sequences.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
1likes39downloads
Dataset Card

Dataset Card for "RO-Offense-Sequences"

Table of Contents

Dataset Description

<!--

  • Paper: News-RO-Offense - A Romanian Offensive Language Dataset and Baseline Models Centered on News Article Comments

-->

Dataset Summary

a novel Romanian language dataset for offensive sequence detection with manually annotated offensive sequences from a local Romanian sports news website (gsp.ro):

Resulting in 4800 annotated messages

Languages

Romanian

Dataset Structure

Data Instances

An example of 'train' looks as follows.

{
  'id': 5,
  'text':'PLACEHOLDER TEXT',
  'offensive_substrings': ['substr1','substr2'],
  'offensive_sequences': [(0,10), (16,20)]
}

Data Fields

  • id: The unique comment ID, corresponding to the ID in RO Offense
  • text: full comment text
  • offensive_substrings: a list of offensive substrings. Can contain duplicates if some offensive substring appears twice
  • offensive_sequences: a list of tuples with (start, end) position of the offensive sequences

Attention: the sequences are computed for \n line sepparator! Git might convert the csv to \r\n.

Data Splits

nametrainvalidatetest
ro4,000400400

Dataset Creation

Curation Rationale

Collecting data for abusive language classification for Romanian Language.

Source Data

Sports News Articles comments

Initial Data Collection and Normalization
Who are the source language producers?

Sports News Article readers

Annotations

Annotation process
Who are the annotators?

Native speakers

Personal and Sensitive Information

The data was public at the time of collection. PII removal has been performed.

Considerations for Using the Data

Social Impact of Dataset

The data definitely contains abusive language. The data could be used to develop and propagate offensive language against every target group involved, i.e. ableism, racism, sexism, ageism, and so on.

Discussion of Biases

Other Known Limitations

Additional Information

Dataset Curators

Licensing Information

This data is available and distributed under Apache-2.0 license

Citation Information

tbd

Contributions