CoolFace
Datasetpublic

BEE-spoke-data/wikipedia-20230901.en-deduped

wikipedia - 20230901.en - deduped purpose: train with less data while maintaining (most) of the quality This is really more of a "high quality diverse sample" rather than "we are trying to remove literal duplicate documents". Source dataset: graelo/wikipedia. configs default command: python -m text_dedup.minhash \ --path $ds_name \ --name $dataset_config \ --split $data_split \ --cache_dir "./cache" \ --output $out_dir \ --column… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/wikipedia-20230901.en-deduped.

sourceHugging Facecc-by-sa-3.0updated 9mo agoView on Hugging Face
6likes1.5kdownloads
settings

This repository belongs to BEE-spoke-data on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namewikipedia-20230901.en-deduped
visibilitypublic
licencecc-by-sa-3.0
gatedno
ownerBEE-spoke-data
Account settings
BEE-spoke-data/wikipedia-20230901.en-deduped · CoolFace