datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
simplequestions-sparqltotext
Dataset Card for SimpleQuestions-SPARQLtoText
Dataset Summary
Special version of SimpleQuestions with SPARQL queries formatted for the SPARQL-to-Text task.
JSON fields
The original version of SimpleQuestions is a raw text file listing triples and the natural language question. A JSON version has been generated and augmented with the following fields:
rdf_subject, rdf_property, rdf_object: triple in the Wikidata format (IDs)
nl_subject, nl_property, nl_object:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/simplequestions-sparqltotext.csqa-sparqltotext
Dataset Card for CSQA-SPARQLtoText
Dataset Summary
CSQA corpus (Complex Sequential Question-Answering, see https://amritasaha1812.github.io/CSQA/) is a large corpus for conversational knowledge-based question answering. The version here is augmented with various fields to make it easier to run specific tasks, especially SPARQL-to-text conversion.
The original data has been post-processing as follows:
Verbalization templates were applied on the answers and their entities… See the full description on the dataset page: https://huggingface.co/datasets/Orange/csqa-sparqltotext.lc_quad2-sparqltotext
Dataset Card for LC-QuAD 2.0 - SPARQLtoText version
Dataset Summary
Special version of LC-QuAD 2.0 for the SPARQL-to-Text task
New field simplified_query
New field is named "simplified_query". It results from applying the following step on the field "query":
Replacing URIs with a simpler format with prefix "resource:", "property:" and "ontology:".
Spacing the delimiters (, {, ., }, ).
Adding diversity to some filters which test a number (contains ( ?var… See the full description on the dataset page: https://huggingface.co/datasets/Orange/lc_quad2-sparqltotext.TimeStress
Dataset Card for Dataset Name
TimeStress is a dataset designed to evaluate the robustness of language models (LMs) to the temporal context of factual knowledge. It enables systematic assessment of whether LMs can correctly associate facts with their valid time periods and distinguish between correct and incorrect temporal contexts at varying granularities (year, month, day).
Dataset Details
Dataset Description
TimeStress consists of over 521,000 natural… See the full description on the dataset page: https://huggingface.co/datasets/Orange/TimeStress.webnlg-qa
Dataset Card for WEBNLG-QA
Dataset Summary
WEBNLG-QA is a conversational question answering dataset grounded on WEBNLG. It consists in a set of question-answering dialogues (follow-up question-answer pairs) based on short paragraphs of text. Each paragraph is associated a knowledge graph (from WEBNLG). The questions are associated with SPARQL queries.
Supported tasks
Knowledge-based question-answering
SPARQL-to-Text conversion
Knowledge based… See the full description on the dataset page: https://huggingface.co/datasets/Orange/webnlg-qa.ORANBench
ORANBench: A Streamlined Benchmark for Assessing LLMs in O-RAN
Overview
ORANBench is a streamlined evaluation dataset derived from ORAN-Bench-13K, designed to efficiently assess Large Language Models (LLMs) in the context of Open Radio Access Networks (O-RAN). This benchmark consists of 1,500 multiple-choice questions, with 500 questions randomly sampled from each of three difficulty levels: easy, intermediate, and difficult.
This dataset is part of the ORANSight-2.0 work… See the full description on the dataset page: https://huggingface.co/datasets/prnshv/ORANBench.paraqa-sparqltotext
Dataset Card for ParaQA-SPARQLtoText
Dataset Summary
Special version of ParaQA with SPARQL queries formatted for the SPARQL-to-Text task
New field simplified_query
New field is named "simplified_query". It results from applying the following step on the field "query":
Replacing URIs with a simpler format with prefix "resource:", "property:" and "ontology:".
Spacing the delimiters (, {, ., }, ).
Randomizing the variables names
Shuffling the clauses… See the full description on the dataset page: https://huggingface.co/datasets/Orange/paraqa-sparqltotext.ORANBench
ORANBench: A Streamlined Benchmark for Assessing LLMs in O-RAN
Overview
ORANBench is a streamlined evaluation dataset derived from ORAN-Bench-13K, designed to efficiently assess Large Language Models (LLMs) in the context of Open Radio Access Networks (O-RAN). This benchmark consists of 1,500 multiple-choice questions, with 500 questions randomly sampled from each of three difficulty levels: easy, intermediate, and difficult.
This dataset is part of the ORANSight-2.0 work… See the full description on the dataset page: https://huggingface.co/datasets/qubol/ORANBench.
