HPLT/HPLT2.0_cleaned
NB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. The Cleaned variant of HPLT Datasets v2.0 This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.
Update README.md
Deprecation warning
Update README.md
Updated plots
Camera-ready plots
Internal evaluations first
specified corpora in the card
Update README.md
Upload 2 files
Update README.md
Upload english-comparison-datasets-by-HPLT.png
Dataset card polished
Update README.md
text metrics description
language table
table with languages and sizes
language codes
Update README.md
tasks and subtasks tags
licence
licence and size attributes
Short description
Delete vie_Latn
Delete tur_Latn
Delete nld_Latn
Delete ell_Grek
Delete ara_Arab in favour of ara_Arab_*
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload dataset (part 00005-of-00006)
Upload dataset (part 00004-of-00006)
Upload dataset (part 00003-of-00006)
Upload dataset (part 00002-of-00006)
Upload dataset (part 00006-of-00007)
Upload dataset (part 00005-of-00007)
Upload dataset (part 00004-of-00007)
Upload dataset (part 00003-of-00007)
Upload dataset (part 00002-of-00007)
Upload dataset (part 00001-of-00007)
Upload dataset (part 00000-of-00007)
Upload dataset (part 00004-of-00005)
Upload dataset (part 00003-of-00005)
Upload dataset (part 00002-of-00005)
Upload dataset (part 00001-of-00005)
Upload dataset (part 00000-of-00005)
Upload dataset (part 00008-of-00009)
Upload dataset (part 00007-of-00009)
Upload dataset (part 00006-of-00009)
Upload dataset (part 00005-of-00009)
Upload dataset (part 00004-of-00009)
