Saibo-creator/bookcorpus_deduplicated
Dataset Card for "bookcorpus_deduplicated" Dataset Summary This is a deduplicated version of the original Book Corpus dataset. The Book Corpus (Zhu et al., 2015), which was used to train popular models such as BERT, has a substantial amount of exact-duplicate documents according to Bandy and Vincent (2021) Bandy and Vincent (2021) find that thousands of books in BookCorpus are duplicated, with only 7,185 unique books out of 11,038 total. Effect of deduplication… See the full description on the dataset page: https://huggingface.co/datasets/Saibo-creator/bookcorpus_deduplicated.
This repository belongs to Saibo-creator on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
