CoolFace
Datasetpublic

KIND-Dataset/KIND

Dataset Summary KIND dataset is a new dilectal data dataset. The dataset was a result of a data marathon competition, where the competitor's goal is to respond to as many prompts as possible in their own dialect, within a fixed time frame with as few errors as possible. For more details, please check the paper The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection Data Fields dialect_code: the label that indicates the specific… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/KIND.

sourceHugging Facecc-by-4.0updated 3y agoView on Hugging Face
0likes13downloads
Dataset Card

Dataset Summary

KIND dataset is a new dilectal data dataset. The dataset was a result of a data marathon competition, where the competitor's goal is to respond to as many prompts as possible in their own dialect, within a fixed time frame with as few errors as possible.

For more details, please check the paper The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection

Data Fields

  • —dialect_code: the label that indicates the specific dialect the text belongs to.
  • —sentenceOriginID: the identifier that references the MSA sentence translated (1000000-2000000), or the reference to link the question to the constructed question dataset (2000000-3000000).
  • —textString: the submitted sentence

Citation Information

@inproceedings{yamani-etal-2024-kind,
    title = "The {KIND} Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection",
    author = "Yamani, Asma  and
      Alziyady, Raghad  and
      AlYami, Reem  and
      Albelali, Salma  and
      Albelali, Leina  and
      Almulhim, Jawharah  and
      Alsulami, Amjad  and
      Alfarraj, Motaz  and
      Al-Zaidy, Rabeah",
    editor = "Falk, Neele  and
      Papi, Sara  and
      Zhang, Mike",
    booktitle = "Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop",
    month = mar,
    year = "2024",
    address = "St. Julian{'}s, Malta",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.eacl-srw.3",
    pages = "32--43",
}