ashtok897/european-hplt-v1
European HPLT v1 A multilingual pretraining corpus of 45,031,396 documents (~50.9B estimated tokens, ~190 GB raw JSONL) across 41 European languages, built from HPLT Monolingual v3 high-quality web crawl data. The corpus spans Germanic, Romance, Slavic, Celtic, Baltic, Finno-Ugric, Greek, and other European language families. Every document has an HPLT WDS quality score of 10 or higher (the top of the quality distribution). Quick Start from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/european-hplt-v1.
3481
Nothing at this path on main. The folder may be empty, or the revision may not exist.
