CoolFace
Datasetpublic

Sadanto3933/commonsense_qa

Usage from datasets import load_dataset dataset=load_dataset( "Sadanto3933/commonsense_qa", split="train", ) # ... Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main… See the full description on the dataset page: https://huggingface.co/datasets/Sadanto3933/commonsense_qa.

sourceHugging Facemitupdated 2y agoView on Hugging Face
1likes12downloads
Dataset Card

Usage

python
from datasets import load_dataset

dataset=load_dataset(
	"Sadanto3933/commonsense_qa",
	split="train",
	)

# ...

Dataset Card for "commonsense_qa"

Dataset Description

  • Homepage: https://www.tau-nlp.org/commonsenseqa
  • Repository: https://github.com/jonathanherzig/commonsenseqa
  • Paper: https://arxiv.org/abs/1811.00937
  • Point of Contact: More Information Needed
  • Size of downloaded dataset files: 4.68 MB
  • Size of the generated dataset: 2.18 MB
  • Total amount of disk used: 6.86 MB

Dataset Summary

CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.

Supported Tasks and Leaderboards

More Information Needed

Languages

The dataset is in English (en).

Dataset Structure

Data Instances

An example of 'train' looks as follows:

json
{'id': '075e483d21c29a511267ef62bedc0461',
 'question': 'The sanctions against the school were a punishing blow, and they seemed to what the efforts the school had made to change?',
 'question_concept': 'punishing',
 'choices': {'label': ['A', 'B', 'C', 'D', 'E'],
  'text': ['ignore', 'enforce', 'authoritarian', 'yell at', 'avoid']},
 'answerKey': 'A'}

Data Fields

The data fields are the same among all splits.

default
  • id (str): Unique ID.
  • question: a string feature.
  • question_concept (str): ConceptNet concept associated to the question.
  • choices: a dictionary feature containing:
  • label: a string feature.
  • text: a string feature.
  • answerKey: a string feature.

Dataset Creation

Licensing Information

The dataset is licensed under the MIT License.

See: https://github.com/jonathanherzig/commonsenseqa/issues/5

Citation Information

@inproceedings{talmor-etal-2019-commonsenseqa,
    title = "{C}ommonsense{QA}: A Question Answering Challenge Targeting Commonsense Knowledge",
    author = "Talmor, Alon  and
      Herzig, Jonathan  and
      Lourie, Nicholas  and
      Berant, Jonathan",
    booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)",
    month = jun,
    year = "2019",
    address = "Minneapolis, Minnesota",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/N19-1421",
    doi = "10.18653/v1/N19-1421",
    pages = "4149--4158",
    archivePrefix = "arXiv",
    eprint        = "1811.00937",
    primaryClass  = "cs",
}