CoolFace
Datasetpublic

datablations/oscar-filter

this is the one where we build the suffix array for 25% Oscar and only deduplicate that part - by deduplication I mean removing any document which has an at least 100-char span overlapping with another document in the 25% chunk. This is very strict and preserves only about 20 million documents, so less then 5% of the full Oscar.

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes2.2kdownloads
Dataset Card

this is the one where we build the suffix array for 25% Oscar and only deduplicate that part - by deduplication I mean removing any document which has an at least 100-char span overlapping with another document in the 25% chunk. This is very strict and preserves only about 20 million documents, so less then 5% of the full Oscar.

datablations/oscar-filter · CoolFace