eduagarcia/CrawlPT_dedup
CrawlPT (deduplicated) CrawlPT is a generic Portuguese corpus extracted from various web pages. This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022). The raw version is also available here. Dataset Details Dataset is composed by three corpora: brWaC, C100-PT, OSCAR-2301. brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites. C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/CrawlPT_dedup.
Update README.md
Full dataset uploaded (part 00005-of-00006)
Full dataset uploaded (part 00004-of-00006)
Full dataset uploaded (part 00003-of-00006)
Full dataset uploaded (part 00002-of-00006)
Full dataset uploaded (part 00001-of-00006)
Full dataset uploaded (part 00000-of-00006)
Update README.md
Update README.md
Update README.md
Update README.md
deduped brwac uploaded
deduped OSCAR-2301 uploaded (part 00003-of-00004)
deduped OSCAR-2301 uploaded (part 00002-of-00004)
deduped OSCAR-2301 uploaded (part 00001-of-00004)
deduped OSCAR-2301 uploaded (part 00000-of-00004)
deduped cc100 uploaded (part 00002-of-00003)
deduped cc100 uploaded (part 00001-of-00003)
deduped cc100 uploaded (part 00000-of-00003)
initial commit
