CoolFace
Datasetpublic

legacy-datasets/hate_offensive

Dataset Card for HateOffensive Dataset Summary Supported Tasks and Leaderboards [More Information Needed] Languages English (en) Dataset Structure Data Instances { "count": 3, "hate_speech_annotation": 0, "offensive_language_annotation": 0, "neither_annotation": 3, "label": 2, # "neither" "tweet": "!!! RT @mayasolovely: As a woman you shouldn't complain about cleaning up your house. & as a man… See the full description on the dataset page: https://huggingface.co/datasets/legacy-datasets/hate_offensive.

sourceHugging Facemitupdated 3y agoView on Hugging Face
8likes164downloads
Dataset Card

Dataset Card for HateOffensive

Table of Contents

Dataset Description

  • Homepage : https://arxiv.org/abs/1905.12516
  • Repository : https://github.com/t-davidson/hate-speech-and-offensive-language
  • Paper : https://arxiv.org/abs/1905.12516
  • Leaderboard :
  • Point of Contact : trd54 at cornell dot edu

Dataset Summary

Supported Tasks and Leaderboards

[More Information Needed]

Languages

English (en)

Dataset Structure

Data Instances

{
"count": 3,
 "hate_speech_annotation": 0,
 "offensive_language_annotation": 0,
 "neither_annotation": 3,
 "label": 2,  # "neither"
 "tweet": "!!! RT @mayasolovely: As a woman you shouldn't complain about cleaning up your house. & as a man you should always take the trash out...")
}

Data Fields

count: (Integer) number of users who coded each tweet (min is 3, sometimes more users coded a tweet when judgments were determined to be unreliable, hatespeechannotation: (Integer) number of users who judged the tweet to be hate speech, offensivelanguageannotation: (Integer) number of users who judged the tweet to be offensive, neither_annotation: (Integer) number of users who judged the tweet to be neither offensive nor non-offensive, label: (Class Label) integer class label for majority of CF users (0: 'hate-speech', 1: 'offensive-language' or 2: 'neither'), tweet: (string)

Data Splits

This dataset is not splitted, only the train split is available.

Dataset Creation

Curation Rationale

[More Information Needed]

Source Data

Initial Data Collection and Normalization

[More Information Needed]

Who are the source language producers?

[More Information Needed]

Annotations

Annotation process

[More Information Needed]

Who are the annotators?

[More Information Needed]

Personal and Sensitive Information

Usernames are not anonymized in the dataset.

Considerations for Using the Data

Social Impact of Dataset

[More Information Needed]

Discussion of Biases

[More Information Needed]

Other Known Limitations

[More Information Needed]

Additional Information

Dataset Curators

[More Information Needed]

Licensing Information

MIT License

Citation Information

@inproceedings{hateoffensive, title = {Automated Hate Speech Detection and the Problem of Offensive Language}, author = {Davidson, Thomas and Warmsley, Dana and Macy, Michael and Weber, Ingmar}, booktitle = {Proceedings of the 11th International AAAI Conference on Web and Social Media}, series = {ICWSM '17}, year = {2017}, location = {Montreal, Canada}, pages = {512-515} }

Contributions

Thanks to @MisbahKhan789 for adding this dataset.