CoolFace
Datasetpublic

Saibo-creator/bookcorpus_deduplicated

Dataset Card for "bookcorpus_deduplicated" Dataset Summary This is a deduplicated version of the original Book Corpus dataset. The Book Corpus (Zhu et al., 2015), which was used to train popular models such as BERT, has a substantial amount of exact-duplicate documents according to Bandy and Vincent (2021) Bandy and Vincent (2021) find that thousands of books in BookCorpus are duplicated, with only 7,185 unique books out of 11,038 total. Effect of deduplication… See the full description on the dataset page: https://huggingface.co/datasets/Saibo-creator/bookcorpus_deduplicated.

sourceHugging Faceupdated 4y agoView on Hugging Face
2likes111downloads

Saibo-creator/bookcorpus_deduplicated · main · files are served by the source, never re-hosted here