helloadhavan/CC-FilteredCorpus
English Cleaned Common Crawl Markdown Dataset An English-focused dataset created from Common Crawl, cleaned and converted to Markdown. The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting. Features English-focused Cleaned and filtered web content HTML converted to Markdown Exact and near-duplicate filtering GPT-2 perplexity filtering Stored as compressed Parquet shards Source The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.
This repository belongs to helloadhavan on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
