data-to-text
mtl-data-to-textmvp-data-to-textcodegen-350M-finetuned-int4-llama_text_to_sql_datasetdataeaze-text2sql-codellama_7b_instruct-clinton_text_to_sql_v1Data-to-text-generation-acceleratemt5-base-finetuned-xsum-data_prep_2021_12_26___t1_7.csv___topic_text_google_mt5_basemt5-base-finetuned-xsum-data_prep_2021_12_26___t404_2980.csv___topic_text_google_mt5_basemt5-base-finetuned-xsum-data_prep_2021_12_26___t22027_162754.csv___topic_text_google_mt5_base
mlb_data_to_textThe MLB dataset for data to text generation contains Major League Baseball games statistics and
their human-written summaries.tool-reasoning-sft-CODING-text_to_terminal_v2-sft-tool-use-agent-data-cleaned-rectified
Text to Terminal, v2 — Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned, combined, and thinking-augmented version of muellerzr/text_to_terminal_v2. It pairs natural language instructions with their corresponding terminal/bash commands, now augmented with explicit <think> reasoning traces that model the step-by-step thought process before producing the final command.The restructuring approach is directly… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-text_to_terminal_v2-sft-tool-use-agent-data-cleaned-rectified.task1728_web_nlg_data_to_text
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.autotrain-data-inanimate-insanity-text-to-animation-video
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Violetmae14/autotrain-data-inanimate-insanity-text-to-animation-video.Data-to-text-GenerationThis is the data-to-text generation datasets collected by TextBox, including:
WebNLG v2.1 (webnlg)
WebNLG v3.0 (webnlg2)
WikiBio (wikibio)
E2E (e2e)
DART (dart)
ToTTo (totto)
ENT-DESC (ent)
AGENDA (agenda)
GenWiki (genwiki)
TEKGEN (tekgen)
LogicNLG (logicnlg)
WikiTableT (wikit)
WEATHERGOV (wg).
The detail and leaderboard of each dataset can be found in TextBox page.
adapt-pre-trained-VL-models-to-text-data-WikipediaThe Wikipedia train data used to train BERT-base baselines and adapt vision-and-language models to text-only tasks in the paper "How to Adapt Pre-trained Vision-and-Language Models to a Text-only Input?".
The data has been created from the "20200501.en" revision of the wikipedia dataset on Huggingface.
