CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01GEM /web_nlgWebNLG is a bi-lingual dataset (English, Russian) of parallel DBpedia triple sets and short texts that cover about 450 different DBpedia properties. The WebNLG data was originally created to promote the development of RDF verbalisers able to generate short text and to handle micro-planning (i.e., sentence segmentation and ordering, referring expression generation, aggregation); the goal of the task is to generate texts starting from 1 to 7 input triples which have entities in common (so the input is actually a connected Knowledge Graph). The dataset contains about 17,000 triple sets and 45,000 crowdsourced texts in English, and 7,000 triples sets and 19,000 crowdsourced texts in Russian. A challenging test set section with entities and/or properties that have not been seen at training time is available.texttable-to-text10K<n<100K4 likes3.4k downloads4y agoHugging Face02webnlg-challenge /web_nlgThe WebNLG challenge consists in mapping data to text. The training data consists of Data/Text pairs where the data is a set of triples extracted from DBpedia and the text is a verbalisation of these triples. For instance, given the 3 DBpedia triples shown in (a), the aim is to generate a text such as (b). a. (John_E_Blaha birthDate 1942_08_26) (John_E_Blaha birthPlace San_Antonio) (John_E_Blaha occupation Fighter_pilot) b. John E Blaha, born in San Antonio on 1942-08-26, worked as a fighter pilot As the example illustrates, the task involves specific NLG subtasks such as sentence segmentation (how to chunk the input data into sentences), lexicalisation (of the DBpedia properties), aggregation (how to avoid repetitions) and surface realisation (how to build a syntactically correct and natural sounding text).tabular-to-text10K<n<100K25 likes868 downloads3y agoHugging Face03webnlg /challenge-2023The WebNLG challenge consists in mapping data to text. The training data consists of Data/Text pairs where the data is a set of triples extracted from DBpedia and the text is a verbalisation of these triples. For instance, given the 3 DBpedia triples shown in (a), the aim is to generate a text such as (b). a. (John_E_Blaha birthDate 1942_08_26) (John_E_Blaha birthPlace San_Antonio) (John_E_Blaha occupation Fighter_pilot) b. John E Blaha, born in San Antonio on 1942-08-26, worked as a fighter pilot As the example illustrates, the task involves specific NLG subtasks such as sentence segmentation (how to chunk the input data into sentences), lexicalisation (of the DBpedia properties), aggregation (how to avoid repetitions) and surface realisation (how to build a syntactically correct and natural sounding text).texttabular-to-text10K<n<100K4 likes395 downloads4y agoHugging Face04Aunderline /webnlgtext1K<n<10K0 likes197 downloads4y agoHugging Face05Aunderline /webnlg_startext1K<n<10K0 likes158 downloads4y agoHugging Face06nishita /webnlg-data2texttext10K<n<100K1 likes124 downloads4y agoHugging Face07Orange /webnlg-qa Dataset Card for WEBNLG-QA Dataset Summary WEBNLG-QA is a conversational question answering dataset grounded on WEBNLG. It consists in a set of question-answering dialogues (follow-up question-answer pairs) based on short paragraphs of text. Each paragraph is associated a knowledge graph (from WEBNLG). The questions are associated with SPARQL queries. Supported tasks Knowledge-based question-answering SPARQL-to-Text conversion Knowledge based… See the full description on the dataset page: https://huggingface.co/datasets/Orange/webnlg-qa.textquestion-answering10K<n<100K1 likes88 downloads3y agoHugging Face08bdsaglam /web_nlg-erx-concattext10K<n<100K0 likes45 downloads2y agoHugging Face09nishita /webnlg_tokenstext1K<n<10K0 likes44 downloads4y agoHugging Face10teven /webnlg_2017_human_evaltabular1K<n<10K0 likes32 downloads4y agoHugging Face11bdsaglam /web_nlg-erxtext10K<n<100K0 likes24 downloads2y agoHugging Face12bdsaglam /web_nlg-erx-concat-chattext10K<n<100K0 likes20 downloads2y agoHugging Face13sazirarrwth99 /web_nlg_devtext1K<n<10K0 likes19 downloads2y agoHugging Face14teven /webnlg_2020_human_evaltabular1K<n<10K0 likes17 downloads4y agoHugging Face15sazirarrwth99 /web_nlg_enriched_refined_new_03text1K<n<10K0 likes17 downloads3y agoHugging Face16sazirarrwth99 /web_nlg_testtext1K<n<10K0 likes16 downloads2y agoHugging Face17sazirarrwth99 /web_nlg_enriched_refined_85_percenttext1K<n<10K0 likes11 downloads2y agoHugging Face18Sachinkelenjaguri /webnlg_Table_to_Texttext10K<n<100K3 likes9 downloads4y agoHugging Face19CATIE-AQ /web_nlg_french Description French translation of the WebNLG dataset.According of the English version dataset card, WebNLG is a bi-lingual dataset (English, Russian) of parallel DBpedia triple sets and short texts that cover about 450 different DBpedia properties. The WebNLG data was originally created to promote the development of RDF verbalisers able to generate short text and to handle micro-planning (i.e., sentence segmentation and ordering, referring expression generation, aggregation); the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/web_nlg_french.texttable-to-text10K<n<100K0 likes9 downloads1y agoHugging Face20sazirarrwth99 /web_nlg_missingtext1K<n<10K0 likes6 downloads3y agoHugging Face21Yomm1927 /WebNLG_KR WebNLG 번역 데이터셋 학습용: ~36k pair 평가용: ~4k pair text10K<n<100K0 likes6 downloads2y agoHugging Face22masterjae /webnlg-kotext10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.