CoolFace
Datasetpublic

rizquuula/commonsense_qa-ID

CommonsenseQA-ID is Indonesian translation version of CommonsenseQA, a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes18downloads
Dataset Card

Dataset Card for "commonsense_qa-ID"

Dataset Description

  • —Homepage: https://github.com/rizquuula/commonsense_qa-ID
  • —Repository: https://github.com/rizquuula/commonsense_qa-ID

Dataset Summary

CommonsenseQA-ID is Indonesian translation version of CommonsenseQA, translated using Google Translation API v2/v3 Basic, all code used for the translation process available in our public repository.

CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see original paper for details.

Languages

The dataset is in Indonesian (id).

Dataset Structure

Data Instances

default
  • —Size of downloaded dataset files: 4.68 MB
  • —Size of the generated dataset: 2.18 MB
  • —Total amount of disk used: 6.86 MB

An example of 'train' looks as follows:

{
  'id': '61fe6e879ff18686d7552425a36344c8',
  'question': 'Sammy ingin pergi ke tempat orang-orang itu berada. Ke mana dia bisa pergi?',
  'question_concept': 'rakyat',
  'choices': {
    'label': ['A', 'B', 'C', 'D', 'E'],
    'text': ['trek balap', 'daerah berpenduduk', 'gurun pasir', 'Apartemen', 'penghalang jalan']
  },
  'answerKey': 'B'
}

Data Fields

The data fields are the same among all splits.

default
  • —id (str): Unique ID.
  • —question: a string feature.
  • —question_concept (str): ConceptNet concept associated to the question.
  • —choices: a dictionary feature containing:
  • —label: a string feature.
  • —text: a string feature.
  • —answerKey: a string feature.

Data Splits

nametrainvalidationtest
default974112211140

Licensing Information

The dataset is licensed under the MIT License.

Citation Information

@inproceedings{talmor-etal-2019-commonsenseqa,
    title = "{C}ommonsense{QA}: A Question Answering Challenge Targeting Commonsense Knowledge",
    author = "Talmor, Alon  and
      Herzig, Jonathan  and
      Lourie, Nicholas  and
      Berant, Jonathan",
    booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)",
    month = jun,
    year = "2019",
    address = "Minneapolis, Minnesota",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/N19-1421",
    doi = "10.18653/v1/N19-1421",
    pages = "4149--4158",
    archivePrefix = "arXiv",
    eprint        = "1811.00937",
    primaryClass  = "cs",
}