datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikilarge-clean
WikiLarge Cleaned
SummaryThis dataset is a cleaned and deduplicated subset of the classic WikiLarge-style sentence pairs (English Wikipedia → Simple English Wikipedia).Starting from the original alignment files (wiki.full.aner.ori.train/valid/test.{src,dst}), we constructed a Hugging Face datasets corpus, applied a set of cheap filters, and removed near-duplicates.
Provenance & License: This is a derivative of Wikipedia / Simple English Wikipedia content under CC BY-SA.The… See the full description on the dataset page: https://huggingface.co/datasets/eilamc14/wikilarge-clean.wikilarge_tswikilarge
