WikiLarge
bart-base-wikilarge-simplificationbart-large-simplification-wikilarge-original-penalizedbart-large-wikilarget5-small-wikilargebart-base-simplification-wikilarge-originalbart-large-simplification-wikilarge-originalt5-small-wikilarge-text-simplificationt5-small-wikilarge-text-simplification-penalty-loss
wikilarge-graded-gpt2toneizerwikilarge-clean
WikiLarge Cleaned
SummaryThis dataset is a cleaned and deduplicated subset of the classic WikiLarge-style sentence pairs (English Wikipedia → Simple English Wikipedia).Starting from the original alignment files (wiki.full.aner.ori.train/valid/test.{src,dst}), we constructed a Hugging Face datasets corpus, applied a set of cheap filters, and removed near-duplicates.
Provenance & License: This is a derivative of Wikipedia / Simple English Wikipedia content under CC BY-SA.The… See the full description on the dataset page: https://huggingface.co/datasets/eilamc14/wikilarge-clean.wikilarge_grade6_alpacawikilarge-text-simplificationwikilarge_grade9_alpacawikilarge
WikiLarge
HuggingFace implementation of the WikiLarge corpus for sentence simplification gathered by Zhang, Xingxing and Lapata, Mirella.
/!\ I am not one of the creators of the dataset, I just needed a HF version of this dataset and uploaded it. I encourage you to read the paper introducing the dataset: Sentence Simplification with Deep Reinforcement Learning (Zhang & Lapata, EMNLP 2017)
Uses
This dataset can be used to train sentence simplification… See the full description on the dataset page: https://huggingface.co/datasets/waboucay/wikilarge.
