CoolFace
Datasetpublic

helloadhavan/CC-FilteredCorpus

English Cleaned Common Crawl Markdown Dataset An English-focused dataset created from Common Crawl, cleaned and converted to Markdown. The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting. Features English-focused Cleaned and filtered web content HTML converted to Markdown Exact and near-duplicate filtering GPT-2 perplexity filtering Stored as compressed Parquet shards Source The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.

sourceHugging Facecc0-1.0updated 1mo agoView on Hugging Face
1likes251downloads
settings

This repository belongs to helloadhavan on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameCC-FilteredCorpus
visibilitypublic
licencecc0-1.0
gatedno
ownerhelloadhavan
Account settings
helloadhavan/CC-FilteredCorpus · CoolFace