CoolFace
Datasetpublic

Orange/csqa-sparqltotext

Dataset Card for CSQA-SPARQLtoText Dataset Summary CSQA corpus (Complex Sequential Question-Answering, see https://amritasaha1812.github.io/CSQA/) is a large corpus for conversational knowledge-based question answering. The version here is augmented with various fields to make it easier to run specific tasks, especially SPARQL-to-text conversion. The original data has been post-processing as follows: Verbalization templates were applied on the answers and their… See the full description on the dataset page: https://huggingface.co/datasets/Orange/csqa-sparqltotext.

sourceHugging Facecc-by-sa-4.0updated 3y agoView on Hugging Face
1likes244downloads
Dataset Card

Dataset Card for CSQA-SPARQLtoText

Table of Contents

Dataset Description

Dataset Summary

CSQA corpus (Complex Sequential Question-Answering, see https://amritasaha1812.github.io/CSQA/) is a large corpus for conversational knowledge-based question answering. The version here is augmented with various fields to make it easier to run specific tasks, especially SPARQL-to-text conversion.

The original data has been post-processing as follows:

  1. 1.Verbalization templates were applied on the answers and their entities were verbalized (replaced by their label in Wikidata)
  1. 1.Questions were parsed using the CARTON algorithm to produce a sequence of action in a specific grammar
  1. 1.Sequence of actions were mapped to SPARQL queries and entities were verbalized (replaced by their label in Wikidata)

Supported tasks

  • Knowledge-based question-answering
  • Text-to-SPARQL conversion
Knowledge based question-answering

Below is an example of dialogue:

  • Q1: Which occupation is the profession of Edmond Yernaux ?
  • A1: politician
  • Q2: Which collectable has that occupation as its principal topic ?
  • A2: Notitia Parliamentaria, An History of the Counties, etc.
SPARQL queries and natural language questions
SQL
SELECT DISTINCT ?x WHERE
{ ?x rdf:type ontology:occupation . resource:Edmond_Yernaux property:occupation ?x }

is equivalent to:

txt
Which occupation is the profession of Edmond Yernaux ?

Languages

  • English

Dataset Structure

The corpus follows the global architecture from the original version of CSQA (https://amritasaha1812.github.io/CSQA/).

There is one directory of the train, dev, and test sets, respectively.

Dialogues are stored in separate directories, 100 dialogues per directory.

Finally, each dialogue is stored in a JSON file as a list of turns.

Types of questions

Comparison of question types compared to related datasets:

[SimpleQuestions](https://huggingface.co/datasets/OrangeInnov/simplequestions-sparqltotext)[ParaQA](https://huggingface.co/datasets/OrangeInnov/paraqa-sparqltotext)[LC-QuAD 2.0](https://huggingface.co/datasets/OrangeInnov/lcquad_2.0-sparqltotext)[CSQA](https://huggingface.co/datasets/OrangeInnov/csqa-sparqltotext)[WebNLQ-QA](https://huggingface.co/datasets/OrangeInnov/webnlg-qa)
Number of triplets in query1
2
More
Logical connector between tripletsConjunction
Disjunction
Exclusion
Topology of the query graphDirect
Sibling
Chain
Mixed
Other
Variable typing in the queryNone
Target variable
Internal variable
Comparisons clausesNone
String
Number
Date
Superlative clausesNo
Yes
Answer typeEntity (open)
Entity (closed)
Number
Boolean
Answer cardinality0 (unanswerable)
1
More
Number of target variables0 (⇒ ASK verb)
1
2
Dialogue contextSelf-sufficient
Coreference
Ellipsis
MeaningMeaningful
Non-sense

Data splits

Text verbalization is only available for a subset of the test set, referred to as challenge set. Other sample only contain dialogues in the form of follow-up sparql queries.

TrainValidationTest
Questions1.5M167K260K
Dialogues152K17K28K
NL question per query1
Characters per query163 (± 100)
Tokens per question10 (± 4)

JSON fields

Each turn of a dialogue contains the following fields:

Original fields
  • ques_type_id: ID corresponding to the question utterance
  • description: Description of type of question
  • relations: ID's of predicates used in the utterance
  • entities_in_utterance: ID's of entities used in the question
  • speaker: The nature of speaker: SYSTEM or USER
  • utterance: The utterance: either the question, clarification or response
  • active_set: A regular expression which identifies the entity set of answer list
  • all_entities: List of ALL entities which constitute the answer of the question
  • question-type: Type of question (broad types used for evaluation as given in the original authors' paper)
  • type_list: List containing entity IDs of all entity parents used in the question
New fields
  • is_spurious: introduced by CARTON,
  • is_incomplete: either the question is self-sufficient (complete) or it relies on information given by the previous turns (incomplete)
  • parsed_active_set:
  • gold_actions: sequence of ACTIONs as returned by CARTON
  • sparql_query: SPARQL query
Verbalized fields

Fields with verbalized in their name are verbalized versions of another fields, ie IDs were replaced by actual words/labels.

Format of the SPARQL queries

  • Clauses are in random order
  • Variables names are represented as random letters. The letters change from one turn to another.
  • Delimiters are spaced

Additional Information

Licensing Information

  • Content from original dataset: CC-BY-SA 4.0
  • New content: CC BY-SA 4.0

Citation Information

This version of the corpus (with SPARQL queries)
bibtex
@inproceedings{lecorve2022sparql2text,
  title={SPARQL-to-Text Question Generation for Knowledge-Based Conversational Applications},
  author={Lecorv\'e, Gw\'enol\'e and Veyret, Morgan and Brabant, Quentin and Rojas-Barahona, Lina M.},
  journal={Proceedings of the Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the International Joint Conference on Natural Language Processing (AACL-IJCNLP)},
  year={2022}
}
Original corpus (CSQA)
bibtex
@InProceedings{saha2018complex,
	title = {Complex {Sequential} {Question} {Answering}: {Towards} {Learning} to {Converse} {Over} {Linked} {Question} {Answer} {Pairs} with a {Knowledge} {Graph}},
	volume = {32},
	issn = {2374-3468},
	url = {https://ojs.aaai.org/index.php/AAAI/article/view/11332},
	booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
	author = {Saha, Amrita and Pahuja, Vardaan and Khapra, Mitesh and Sankaranarayanan, Karthik and Chandar, Sarath},
	month = apr,
	year = {2018}
}
CARTON
bibtex
@InProceedings{plepi2021context,
	author="Plepi, Joan and Kacupaj, Endri and Singh, Kuldeep and Thakkar, Harsh and Lehmann, Jens",
	editor="Verborgh, Ruben and Hose, Katja and Paulheim, Heiko and Champin, Pierre-Antoine and Maleshkova, Maria and Corcho, Oscar and Ristoski, Petar and Alam, Mehwish",
	title="Context Transformer with Stacked Pointer Networks for Conversational Question Answering over Knowledge Graphs",
	booktitle="Proceedings of The Semantic Web",
	year="2021",
	publisher="Springer International Publishing",
	pages="356--371",
	isbn="978-3-030-77385-4"
}