datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
web_nlgWebNLG is a bi-lingual dataset (English, Russian) of parallel DBpedia triple sets
and short texts that cover about 450 different DBpedia properties. The WebNLG data
was originally created to promote the development of RDF verbalisers able to
generate short text and to handle micro-planning (i.e., sentence segmentation and
ordering, referring expression generation, aggregation); the goal of the task is
to generate texts starting from 1 to 7 input triples which have entities in common
(so the input is actually a connected Knowledge Graph). The dataset contains about
17,000 triple sets and 45,000 crowdsourced texts in English, and 7,000 triples sets
and 19,000 crowdsourced texts in Russian. A challenging test set section with
entities and/or properties that have not been seen at training time is available.NLG-Machine-Translation
SEA Machine Translation
SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.NLG-Abstractive-Summarization
SEA Abstractive Summarization
SEA Abstractive Summarization evaluates a model's ability to read a document, identify the key points within, and summarize them into a coherent and fluent text while paraphrasing the document. It is sampled from XL-Sum for Indonesian, Tamil, Thai, and Vietnamese.
Supported Tasks and Leaderboards
SEA Abstractive Summarization is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Abstractive-Summarization.e2e_nlgThe E2E dataset is designed for a limited-domain data-to-text task --
generation of restaurant descriptions/recommendations based on up to 8 different
attributes (name, area, price range etc.).nl_gameable_programmatic_graderstask1728_web_nlg_data_to_text
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.recipe_nlgtask957_e2e_nlg_text_generation_generate
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task957_e2e_nlg_text_generation_generate
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task957_e2e_nlg_text_generation_generate.e2e_nlg_cleaned_promptsourcerecipe-nlg-llama2
Dataset Card for "recipe-nlg-llama2"
More Information needed
nlgraph
Dataset Card for "nlgraph"
@article{wang2023can,
title={Can Language Models Solve Graph Problems in Natural Language?},
author={Wang, Heng and Feng, Shangbin and He, Tianxing and Tan, Zhaoxuan and Han, Xiaochuang and Tsvetkov, Yulia},
journal={arXiv preprint arXiv:2305.10037},
year={2023}
}
recipe-nlg-50kweb_nlg-erx-concatmsmarco-nlgen
Dataset Card for MSMARCO - Natural Language Generation Task
Dataset Summary
The original focus of MSMARCO was to provide a corpus for training and testing systems which given a real domain user query systems would then provide the most likley candidate answer and do so in language which was natural and conversational. All questions have been generated from real anonymized Bing user queries which grounds the dataset in a real world problem and can provide researchers real… See the full description on the dataset page: https://huggingface.co/datasets/din0s/msmarco-nlgen.e2e-nlg-chatmlnlg_itenriched_web_nlg_en_promptsourcenlg_denlg_enweb_nlg-erxweb_nlg_devweb_nlg-erx-concat-chatweb_nlg_enriched_refined_new_03flan_combined_task1728_web_nlg_data_to_textweb_nlg_testrankme-nlg-acceptability@inproceedings{novikova-etal-2018-rankme,
title = "RankME: Reliable Human Ratings for Natural Language Generation",
author = "Novikova, Jekaterina and
Duvsek, Ondvrej and
Rieser, Verena",
booktitle = "Proceedings of the NAACL2018",
month = jun,
year = "2018",
address = "New Orleans, Louisiana",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/N18-2012",
doi = "10.18653/v1/N18-2012",
pages = "72--78"… See the full description on the dataset page: https://huggingface.co/datasets/metaeval/rankme-nlg-acceptability.nlg_mix_en_de_itrecipe-nlg-lite-llama-2
Dataset Card for "recipe-nlg-lite-llama-2"
More Information needed
web_nlg_enriched_refined_85_percentTripletDollyQA-3k-Gemma-Nlg
