CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01GEM /web_nlgWebNLG is a bi-lingual dataset (English, Russian) of parallel DBpedia triple sets and short texts that cover about 450 different DBpedia properties. The WebNLG data was originally created to promote the development of RDF verbalisers able to generate short text and to handle micro-planning (i.e., sentence segmentation and ordering, referring expression generation, aggregation); the goal of the task is to generate texts starting from 1 to 7 input triples which have entities in common (so the input is actually a connected Knowledge Graph). The dataset contains about 17,000 triple sets and 45,000 crowdsourced texts in English, and 7,000 triples sets and 19,000 crowdsourced texts in Russian. A challenging test set section with entities and/or properties that have not been seen at training time is available.texttable-to-text10K<n<100K4 likes3.2k downloads4y agoHugging Face02aisingapore /NLG-Machine-Translationgated SEA Machine Translation SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese. Supported Tasks and Leaderboards SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.texttext-generation10K<n<100K2 likes2.9k downloads9mo agoHugging Face03aisingapore /NLG-Abstractive-Summarizationgated SEA Abstractive Summarization SEA Abstractive Summarization evaluates a model's ability to read a document, identify the key points within, and summarize them into a coherent and fluent text while paraphrasing the document. It is sampled from XL-Sum for Indonesian, Tamil, Thai, and Vietnamese. Supported Tasks and Leaderboards SEA Abstractive Summarization is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Abstractive-Summarization.texttext-generationn<1K0 likes2.3k downloads9mo agoHugging Face04GEM /e2e_nlgThe E2E dataset is designed for a limited-domain data-to-text task -- generation of restaurant descriptions/recommendations based on up to 8 different attributes (name, area, price range etc.).texttable-to-text10K<n<100K2 likes726 downloads4y agoHugging Face05oliverdk /nl_gameable_programmatic_graderstextn<1K0 likes206 downloads3mo agoHugging Face06Lots-of-LoRAs /task1728_web_nlg_data_to_text Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.texttext-generation1K<n<10K0 likes178 downloads2y agoHugging Face07Zappandy /recipe_nlgtext100K<n<1M4 likes148 downloads4y agoHugging Face08Lots-of-LoRAs /task957_e2e_nlg_text_generation_generate Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task957_e2e_nlg_text_generation_generate Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task957_e2e_nlg_text_generation_generate.texttext-generation1K<n<10K0 likes116 downloads2y agoHugging Face09marcov /e2e_nlg_cleaned_promptsourcetext100K<n<1M0 likes102 downloads2y agoHugging Face10skadewdl3 /recipe-nlg-llama2 Dataset Card for "recipe-nlg-llama2" More Information needed text1M<n<10M3 likes80 downloads3y agoHugging Face11tasksource /nlgraph Dataset Card for "nlgraph" @article{wang2023can, title={Can Language Models Solve Graph Problems in Natural Language?}, author={Wang, Heng and Feng, Shangbin and He, Tianxing and Tan, Zhaoxuan and Han, Xiaochuang and Tsvetkov, Yulia}, journal={arXiv preprint arXiv:2305.10037}, year={2023} } text1K<n<10K5 likes70 downloads3y agoHugging Face12EmTpro01 /recipe-nlg-50ktext10K<n<100K0 likes52 downloads2y agoHugging Face13bdsaglam /web_nlg-erx-concattext10K<n<100K0 likes47 downloads2y agoHugging Face14din0s /msmarco-nlgen Dataset Card for MSMARCO - Natural Language Generation Task Dataset Summary The original focus of MSMARCO was to provide a corpus for training and testing systems which given a real domain user query systems would then provide the most likley candidate answer and do so in language which was natural and conversational. All questions have been generated from real anonymized Bing user queries which grounds the dataset in a real world problem and can provide researchers real… See the full description on the dataset page: https://huggingface.co/datasets/din0s/msmarco-nlgen.textquestion-answering100K<n<1M6 likes46 downloads4y agoHugging Face15gomind /e2e-nlg-chatmltext10K<n<100K0 likes42 downloads1mo agoHugging Face16Jiahuan /nlg_ittext1K<n<10K0 likes33 downloads3y agoHugging Face17marcov /enriched_web_nlg_en_promptsourcetext10K<n<100K0 likes32 downloads2y agoHugging Face18Jiahuan /nlg_detext1K<n<10K0 likes31 downloads3y agoHugging Face19Jiahuan /nlg_entext1K<n<10K0 likes28 downloads3y agoHugging Face20bdsaglam /web_nlg-erxtext10K<n<100K0 likes24 downloads2y agoHugging Face21sazirarrwth99 /web_nlg_devtext1K<n<10K0 likes19 downloads2y agoHugging Face22bdsaglam /web_nlg-erx-concat-chattext10K<n<100K0 likes19 downloads2y agoHugging Face23sazirarrwth99 /web_nlg_enriched_refined_new_03text1K<n<10K0 likes17 downloads2y agoHugging Face24supergoose /flan_combined_task1728_web_nlg_data_to_texttext10K<n<100K0 likes17 downloads2y agoHugging Face25sazirarrwth99 /web_nlg_testtext1K<n<10K0 likes16 downloads2y agoHugging Face26metaeval /rankme-nlg-acceptability@inproceedings{novikova-etal-2018-rankme, title = "RankME: Reliable Human Ratings for Natural Language Generation", author = "Novikova, Jekaterina and Duvsek, Ondvrej and Rieser, Verena", booktitle = "Proceedings of the NAACL2018", month = jun, year = "2018", address = "New Orleans, Louisiana", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/N18-2012", doi = "10.18653/v1/N18-2012", pages = "72--78"… See the full description on the dataset page: https://huggingface.co/datasets/metaeval/rankme-nlg-acceptability.tabulartext-classification1K<n<10K0 likes15 downloads4y agoHugging Face27Jiahuan /nlg_mix_en_de_ittext10K<n<100K0 likes12 downloads3y agoHugging Face28skadewdl3 /recipe-nlg-lite-llama-2 Dataset Card for "recipe-nlg-lite-llama-2" More Information needed text1K<n<10K1 likes11 downloads3y agoHugging Face29sazirarrwth99 /web_nlg_enriched_refined_85_percenttext1K<n<10K0 likes11 downloads2y agoHugging Face30rPucs /TripletDollyQA-3k-Gemma-Nlgtext1K<n<10K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.