CoolFace
Datasetpublic

timonziegenbein/appropriateness-corpus

The Appropriateness Corpus The Appropriateness Corpus is a collection of 2191 arguments annotated for appropriateness and its 14 subdimensions derived in the paper Modeling Appropriate Language in Argumentation published at the ACL2023. Dataset Description What does Appropriateness mean? An argument “has an appropriate style if the used language supports the creation of credibility and emotions as well as if it is proportional to the issue.”… See the full description on the dataset page: https://huggingface.co/datasets/timonziegenbein/appropriateness-corpus.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes44downloads
Dataset Card

The Appropriateness Corpus

<!-- Provide a quick summary of the dataset. -->

The Appropriateness Corpus is a collection of 2191 arguments annotated for appropriateness and its 14 subdimensions derived in the paper Modeling Appropriate Language in Argumentation published at the ACL2023.

Dataset Description

<!-- Provide a longer summary of what this dataset is. -->

What does Appropriateness mean?

An argument “has an appropriate style if the used language supports the creation of credibility and emotions as well as if it is proportional to the issue.” Their annotation guidelines further suggest that “the choice of words and the grammatical complexity should [...] appear suitable for the topic discussed within the given setting [...], matching the way credibility and emotions are created [...]”.

Wachsmuth et al. (2017)

What makes an Argument (In)appropriate?

<img src="https://raw.githubusercontent.com/timonziegenbein/appropriateness-corpus/main/annotation-guidelines/appropriateness-taxonomy-vertical.svg">

Toxic Emotions (TE): An argument has toxic emotions if the emotions appealed to are deceptive or their intensities do not provide room for critical evaluation of the issue by the reader.

  • Excessive Intensity (EI): The emotions appealed to by an argument are unnecessarily strong for the discussed issue.
  • Emotional Deception (ED): The emotions appealed to are used as deceptive tricks to win, derail, or end the discussion.

Missing Commitment (MC): An argument is missing commitment if the issue is not taken seriously or openness other’s arguments is absent.

  • Missing Seriousness (MS): The argument is either trolling others by suggesting (explicitly or implicitly) that the issue is not worthy of being discussed or does not contribute meaningfully to the discussion.
  • Missing Openness (MO): The argument displays an unwillingness to consider arguments with opposing viewpoints and does not assess the arguments on their merits but simply rejects them out of hand.

Missing Intelligibility (MI): An argument is not intelligible if its meaning is unclear or irrelevant to the issue or if its reasoning is not understandable.

  • Unclear Meaning (UM): The argument’s content is vague, ambiguous, or implicit, such that it remains unclear what is being said about the issue (it could also be an unrelated issue).
  • Missing Relevance (MR): The argument does not discuss the issue, but derails the discussion implicitly towards a related issue or shifts completely towards a different issue.
  • Confusing Reasoning (CR): The argument’s components (claims and premises) seem not to be connected logically.

Other Reasons (OR): An argument is inappropriate if it contains severe orthographic errors or for reasons not covered by any other dimension.

  • Detrimental Orthography (DO): The argument has serious spelling and/or grammatical errors, negatively affecting its readability.
  • Reason Unclassified (RU): There are any other reasons than those above for why the argument should be considered inappropriate.

Dataset Sources

<!-- Provide the basic links for the dataset. -->

Dataset Structure

<!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. -->

The dataset columns mostly present the different appropriateness flaws explained above; if the value in a column is 1, the respective flaw was annotated to be present by at least one of the annotators. Otherwise, the value will be 0.

Citation

<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. -->

If you are interested in using the corpus, please cite the following paper: Modeling Appropriate Language in Argumentation (Ziegenbein et al., ACL 2023)