datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikilarge-graded-gpt2toneizerwikilarge-clean
WikiLarge Cleaned
SummaryThis dataset is a cleaned and deduplicated subset of the classic WikiLarge-style sentence pairs (English Wikipedia → Simple English Wikipedia).Starting from the original alignment files (wiki.full.aner.ori.train/valid/test.{src,dst}), we constructed a Hugging Face datasets corpus, applied a set of cheap filters, and removed near-duplicates.
Provenance & License: This is a derivative of Wikipedia / Simple English Wikipedia content under CC BY-SA.The… See the full description on the dataset page: https://huggingface.co/datasets/eilamc14/wikilarge-clean.wikilarge_grade6_alpacawikilarge-text-simplificationwikilarge_grade9_alpacawikilarge
WikiLarge
HuggingFace implementation of the WikiLarge corpus for sentence simplification gathered by Zhang, Xingxing and Lapata, Mirella.
/!\ I am not one of the creators of the dataset, I just needed a HF version of this dataset and uploaded it. I encourage you to read the paper introducing the dataset: Sentence Simplification with Deep Reinforcement Learning (Zhang & Lapata, EMNLP 2017)
Uses
This dataset can be used to train sentence simplification… See the full description on the dataset page: https://huggingface.co/datasets/waboucay/wikilarge.wikilarge_alpaca_jsonwikilarge-text-simplificationwikilarge_grade5_alpacagraded_wikilargewikilarge_grade7_alpacawikilarge_grade4_alpacawikilarge_grade3_alpacawikilarge_tswikilarge_grade10_alpacawikilargewikilarge_grade2_alpacawikilarge_grade6_8_alpacawikilarge_grade8_alpacaWikiLarge_ori_splitwisewikilargeimgwikilarge_grade2_12_alpacawikilarge_baseline_alpacawikilarge
WikiLarge
HuggingFace implementation of the WikiLarge corpus for sentence simplification gathered by Zhang, Xingxing and Lapata, Mirella.
/!\ I am not one of the creators of the dataset, I just needed a HF version of this dataset and uploaded it. I encourage you to read the paper introducing the dataset: Sentence Simplification with Deep Reinforcement Learning (Zhang & Lapata, EMNLP 2017)
Uses
This dataset can be used to train sentence… See the full description on the dataset page: https://huggingface.co/datasets/navii23/wikilarge.wikilarge_grade10_12_alpacawikilarge_grade12_alpacawikilarge_grade3_11_alpacawikilarge_grade11_alpacawikilarge_grade4_10_alpacawikilarge_grade2_4_alpaca
