CoolFace
Datasetpublic

fbougares/simple_questions_v2

SimpleQuestions is a dataset for simple QA, which consists of a total of 108,442 questions written in natural language by human English-speaking annotators each paired with a corresponding fact, formatted as (subject, relationship, object), that provides the answer but also a complete explanation. Fast have been extracted from the Knowledge Base Freebase (freebase.com). We randomly shuffle these questions and use 70% of them (75910) as training set, 10% as validation set (10845), and the remaining 20% as test set.

sourceHugging Facecc-by-3.0updated 3y agoView on Hugging Face
3likes359downloads
Dataset Card

Dataset Card for SimpleQuestions

Table of Contents

Dataset Description

  • Homepage: https://research.fb.com/downloads/babi/
  • Repository: https://github.com/fbougares/TSAC
  • Paper: https://research.fb.com/publications/large-scale-simple-question-answering-with-memory-networks/
  • Leaderboard: [If the dataset supports an active leaderboard, add link here]()
  • Point of Contact: Antoine Borde Nicolas Usunie Sumit Chopra, Jason Weston

Dataset Summary

[More Information Needed]

Supported Tasks and Leaderboards

[More Information Needed]

Languages

[More Information Needed]

Dataset Structure

Data Instances

Here are some examples of questions and facts:

  • What American cartoonist is the creator of Andy Lippincott? Fact: (andylippincott, charactercreatedby, garrytrudeau)
  • Which forest is Fires Creek in? Fact: (firescreek, containedby, nantahalanational_forest)
  • What does Jimmy Neutron do? Fact: (jimmyneutron, fictionalcharacter_occupation, inventor)
  • What dietary restriction is incompatible with kimchi? Fact: (kimchi, incompatiblewithdietary_restrictions, veganism)

Data Fields

[More Information Needed]

Data Splits

[More Information Needed]

Dataset Creation

Curation Rationale

[More Information Needed]

Source Data

[More Information Needed]

Initial Data Collection and Normalization

[More Information Needed]

Who are the source language producers?

[More Information Needed]

Annotations

[More Information Needed]

Annotation process

[More Information Needed]

Who are the annotators?

[More Information Needed]

Personal and Sensitive Information

[More Information Needed]

Considerations for Using the Data

Social Impact of Dataset

[More Information Needed]

Discussion of Biases

[More Information Needed]

Other Known Limitations

[More Information Needed]

Additional Information

Dataset Curators

[More Information Needed]

Licensing Information

[More Information Needed]

Citation Information

[More Information Needed]

Contributions

Thanks to @abhishekkrthakur for adding this dataset.