CoolFace
Datasetpublic

LeMoussel/fra-hplt

French HPLT A French pretraining corpus of 720 405 documents built from HPLT Monolingual v3 high-quality web crawl data. The corpus is streaming-ready and includes per-document metadata: language confidence, URL, quality score, character/word counts, and per-register scores. Every document has an HPLT WDS quality score of 10 or higher (the top of the quality distribution). Quick Start from datasets import load_dataset # Streaming ds =… See the full description on the dataset page: https://huggingface.co/datasets/LeMoussel/fra-hplt.

sourceHugging Facecc0-1.0updated 3mo agoView on Hugging Face
1likes90downloads

LeMoussel/fra-hplt · main · files are served by the source, never re-hosted here