CoolFace
Datasetpublic

clarin-knext/cst_directed_datasets

Data for Polish CST task annotated manually

sourceHugging Faceccupdated 3y agoView on Hugging Face
0likes32downloads
Dataset Card

Cross Document Structure Theory datasets

Table of Contents

Dataset Description

  • Homepage:
  • Repository:
  • Paper:
  • Point of Contact: arkadiusz.janz@pwr.edu.pl

Dataset Summary

Collection of sentence pairs annotated for the Cross-Document Structure Theory (CST) task. It consists of 4 distinct datasets:

  • WUT
  • PODCAST
  • SNLI
  • WNLI

Supported Tasks and Leaderboards

[More Information Needed]

Languages

Polish language, PL English language, EN

Dataset Structure

Data Instances

Data are structured in JSONL format, each single text sample is divided by sentence.

{
  "id": 0, 
  "sentence_1": "According to the television coverage of the royal wedding, the bride looked beautiful.", 
  "sentence_2": "The bride looked exceptionally beautiful.", 
  "cst": [
    {
      "supertype": "overlap", 
      "type": "attribution", 
      "knowledge": "text"
    }, 
    {
      "supertype": "overlap", 
      "type": "elaboration", 
      "knowledge": "text"
    }
  ]
}

Data Fields

Description of json keys:

  • id: identifier of the sentence pair
  • sentence_1: first sentence
  • sentence_2: second sentence
  • cst: list of CST labels
  • supertype: relation supertype
  • type: specific relation type
  • knowledge: premise knowledge source

Data Splits

We do not specify an exact data split for training and evaluation.

Dataset Creation

Curation Rationale

[More Information Needed]

Source Data

Initial Data Collection, Normalization and Post-processing

[More Information Needed]

Annotations

[image]

Annotation process

[More Information Needed]

Who are the annotators?
  • professional linguists (mention all people involved)

Personal and Sensitive Information

The datasets do not contain any personal or sensitive information.

Considerations for Using the Data

Discussion of Biases

[More Information Needed]

Other Known Limitations

[More Information Needed]

Additional Information

Dataset Curators

Arkadiusz Janz (arkadiusz.janz@pwr.edu.pl)

Licensing Information

WUT CC-BY 2.5 PODCAST CC-BY-SA 4.0 SNLI CC-BY-SA 4.0 WNLI CC-BY-SA 4.0

Citation Information

[More Information Needed]