CoolFace
Datasetpublic

UppsalaNLP/swedish-causality-ranking

Swedish Causality Ranking Dataset Causality ranking dataset for Swedish text, comparing sentence pairs on how well they match a causal prompt. Dataset Description This dataset contains pairs of Swedish sentences annotated on a 6-point scale for their causal relevance to a given prompt containing a cause or effect. Fields prompt: Query containing a cause or effect (e.g., "Verkan: växthuseffekt") sentence_1_left_context: Context before sentence 1… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-causality-ranking.

sourceHugging Facecc-by-4.0updated 8mo agoView on Hugging Face
0likes18downloads
Dataset Card

Swedish Causality Ranking Dataset

Causality ranking dataset for Swedish text, comparing sentence pairs on how well they match a causal prompt.

Dataset Description

This dataset contains pairs of Swedish sentences annotated on a 6-point scale for their causal relevance to a given prompt containing a cause or effect.

Fields

  • prompt: Query containing a cause or effect (e.g., "Verkan: växthuseffekt")
  • sentence_1_left_context: Context before sentence 1
  • sentence_1_target: First target sentence
  • sentence_1_right_context: Context after sentence 1
  • sentence_2_left_context: Context before sentence 2
  • sentence_2_target: Second target sentence
  • sentence_2_right_context: Context after sentence 2
  • annotation: Ranking score (1-6 scale)

Annotation Scale

The 6-point scale compares which sentence better expresses a causal relation matching the prompt:

  • Lower scores: Sentence 1 is more relevant
  • Higher scores: Sentence 2 is more relevant
  • Middle scores: Both sentences are similarly relevant

Usage

python
from datasets import load_dataset

dataset = load_dataset("UppsalaNLP/swedish-causality-ranking")

# Access the data
data = dataset["train"]

# Example
print(data[0]["prompt"])
print(data[0]["sentence_1_target"])
print(data[0]["sentence_2_target"])
print(data[0]["annotation"])

Source

Text extracted from the SOU-corpus (Swedish Government Official Reports).

Citation

bibtex
@inproceedings{durlich-etal-2022-cause,
    title = "Cause and Effect in Governmental Reports: Two Data Sets for Causality Detection in Swedish",
    author = "D{\"u}rlich, Luise and Reimann, Sebastian and Finnveden, Gustav and Nivre, Joakim and Stymne, Sara",
    booktitle = "Proceedings of the First Workshop on Natural Language Processing for Political Sciences",
    month = jun,
    year = "2022",
    address = "Marseilles, France"
}

License

This dataset is licensed under CC BY 4.0.