LeMoussel/fra-hplt
French HPLT A French pretraining corpus of 720 405 documents built from HPLT Monolingual v3 high-quality web crawl data. The corpus is streaming-ready and includes per-document metadata: language confidence, URL, quality score, character/word counts, and per-register scores. Every document has an HPLT WDS quality score of 10 or higher (the top of the quality distribution). Quick Start from datasets import load_dataset # Streaming ds =… See the full description on the dataset page: https://huggingface.co/datasets/LeMoussel/fra-hplt.
190
