CoolFace
Datasetpublic

MichaelR207/high-quality-cc-21b

high_quality A high-quality English web text corpus extracted from Common Crawl WARC files using an LLM-based extraction and quality pipeline. Dataset Summary high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl WARC records are passed through an LLM-based extractor that strips boilerplate and recovers the main content, then filtered to retain only documents in the "high_quality" band, deduplicated (exact + fuzzy), and… See the full description on the dataset page: https://huggingface.co/datasets/MichaelR207/high-quality-cc-21b.

sourceHugging Faceodc-byupdated 3mo agoView on Hugging Face
0likes321downloads
settings

This repository belongs to MichaelR207 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namehigh-quality-cc-21b
visibilitypublic
licenceodc-by
gatedno
ownerMichaelR207
Account settings
MichaelR207/high-quality-cc-21b · CoolFace