Orange/lc_quad2-sparqltotext
Dataset Card for LC-QuAD 2.0 - SPARQLtoText version Dataset Summary Special version of LC-QuAD 2.0 for the SPARQL-to-Text task New field simplified_query New field is named "simplified_query". It results from applying the following step on the field "query": Replacing URIs with a simpler format with prefix "resource:", "property:" and "ontology:". Spacing the delimiters (, {, ., }, ). Adding diversity to some filters which test a number (contains… See the full description on the dataset page: https://huggingface.co/datasets/Orange/lc_quad2-sparqltotext.
Dataset Card for LC-QuAD 2.0 - SPARQLtoText version
Table of Contents
- Dataset Card for LC-QuAD 2.0 - SPARQLtoText version
- Table of Contents
- Dataset Description
- Dataset Summary
- New field `simplified_query`
- New split "valid"
- Supported tasks
- Languages
- Dataset Structure
- Types of questions
- Data splits
- Additional information
- Related datasets
- Licencing information
- Citation information
- This version of the corpus (with normalized SPARQL queries)
- Original version
Dataset Description
- Paper: SPARQL-to-Text Question Generation for Knowledge-Based Conversational Applications (AACL-IJCNLP 2022)
- Point of Contact: Gwénolé Lecorvé
Dataset Summary
Special version of LC-QuAD 2.0 for the SPARQL-to-Text task
New field simplified_query
New field is named "simplified_query". It results from applying the following step on the field "query":
- Replacing URIs with a simpler format with prefix "resource:", "property:" and "ontology:".
- Spacing the delimiters
(,{,.,},).
- Adding diversity to some filters which test a number (
contains ( ?var, 'number' )can becomecontains ?var = number
- Randomizing the variables names
- Shuffling the clauses
New split "valid"
A validation set was randonly extracted from the test set to represent 10% of the whole dataset.
Supported tasks
- Knowledge-based question-answering
- Text-to-SPARQL conversion
- SPARQL-to-Text conversion
Languages
- English
Dataset Structure
The corpus follows the global architecture from the original version of CSQA (https://amritasaha1812.github.io/CSQA/).
There is one directory of the train, dev, and test sets, respectively.
Dialogues are stored in separate directories, 100 dialogues per directory.
Finally, each dialogue is stored in a JSON file as a list of turns.
Types of questions
Comparison of question types compared to related datasets:
Data splits
Text verbalization is only available for a subset of the test set, referred to as challenge set. Other sample only contain dialogues in the form of follow-up sparql queries.
Additional information
Related datasets
This corpus is part of a set of 5 datasets released for SPARQL-to-Text generation, namely:
- Non conversational datasets
- SimpleQuestions (from https://github.com/askplatypus/wikidata-simplequestions)
- ParaQA (from https://github.com/barshana-banerjee/ParaQA)
- LC-QuAD 2.0 (from http://lc-quad.sda.tech/)
- Conversational datasets
- CSQA (from https://amritasaha1812.github.io/CSQA/)
- WebNLQ-QA (derived from https://gitlab.com/shimorina/webnlg-dataset/-/tree/master/release_v3.0)
Licencing information
- Content from original dataset: CC-BY 3.0
- New content: CC BY-SA 4.0
Citation information
This version of the corpus (with normalized SPARQL queries)
@inproceedings{lecorve2022sparql2text,
title={SPARQL-to-Text Question Generation for Knowledge-Based Conversational Applications},
author={Lecorv\'e, Gw\'enol\'e and Veyret, Morgan and Brabant, Quentin and Rojas-Barahona, Lina M.},
journal={Proceedings of the Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the International Joint Conference on Natural Language Processing (AACL-IJCNLP)},
year={2022}
}Original version
@inproceedings{dubey2017lc2,
title={LC-QuAD 2.0: A Large Dataset for Complex Question Answering over Wikidata and DBpedia},
author={Dubey, Mohnish and Banerjee, Debayan and Abdelkawi, Abdelrahman and Lehmann, Jens},
booktitle={Proceedings of the 18th International Semantic Web Conference (ISWC)},
year={2019},
organization={Springer}
}