CoolFace
14 results

WikiLarge

williamplacroix /wikilarge-graded-gpt2toneizer100K<n<1M0 likes475 downloads2y agoHugging Faceeilamc14 /wikilarge-clean WikiLarge Cleaned SummaryThis dataset is a cleaned and deduplicated subset of the classic WikiLarge-style sentence pairs (English Wikipedia → Simple English Wikipedia).Starting from the original alignment files (wiki.full.aner.ori.train/valid/test.{src,dst}), we constructed a Hugging Face datasets corpus, applied a set of cheap filters, and removed near-duplicates. Provenance & License: This is a derivative of Wikipedia / Simple English Wikipedia content under CC BY-SA.The… See the full description on the dataset page: https://huggingface.co/datasets/eilamc14/wikilarge-clean.text100K<n<1M0 likes132 downloads8mo agoHugging Facewilliamplacroix /wikilarge_grade6_alpacatext10K<n<100K0 likes91 downloads1y agoHugging Facebogdancazan /wikilarge-text-simplificationtext100K<n<1M5 likes90 downloads3y agoHugging Facewilliamplacroix /wikilarge_grade9_alpacatext10K<n<100K0 likes79 downloads1y agoHugging Facewaboucay /wikilarge WikiLarge HuggingFace implementation of the WikiLarge corpus for sentence simplification gathered by Zhang, Xingxing and Lapata, Mirella. /!\ I am not one of the creators of the dataset, I just needed a HF version of this dataset and uploaded it. I encourage you to read the paper introducing the dataset: Sentence Simplification with Deep Reinforcement Learning (Zhang & Lapata, EMNLP 2017) Uses This dataset can be used to train sentence simplification… See the full description on the dataset page: https://huggingface.co/datasets/waboucay/wikilarge.2 likes72 downloads2y agoHugging Face