CoolFace
Datasetpublic

ashtok897/european-hplt-v1

European HPLT v1 A multilingual pretraining corpus of 45,031,396 documents (~50.9B estimated tokens, ~190 GB raw JSONL) across 41 European languages, built from HPLT Monolingual v3 high-quality web crawl data. The corpus spans Germanic, Romance, Slavic, Celtic, Baltic, Finno-Ugric, Greek, and other European language families. Every document has an HPLT WDS quality score of 10 or higher (the top of the quality distribution). Quick Start from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/european-hplt-v1.

sourceHugging Facecc0-1.0updated 4mo agoView on Hugging Face
3likes481downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

ashtok897/european-hplt-v1 · main · files are served by the source, never re-hosted here