CoolFace
Datasetpublic

anandjh8/common-crawl-english-filtered

🧠 FineWeb-English-Filtered 📘 Dataset Summary FineWeb-English-Filtered is a large-scale, cleaned, English-only text dataset derived from Common Crawl’s WET archives.It contains 940 million documents of publicly available web text, converted into Apache Parquet format with a consistent schema for fast and efficient data loading. The dataset was generated using a custom AWS Glue pipeline that processed, filtered, and merged .wet files across multiple terabytes of… See the full description on the dataset page: https://huggingface.co/datasets/anandjh8/common-crawl-english-filtered.

sourceHugging Facecc-by-sa-4.0updated 4mo agoView on Hugging Face
2likes324downloads
5 commits on main
fc2cede4mo ago

Update README.md

anandjh8
d0762501y ago

Create README.md

anandjh8
4cbedcf1y ago

Upload folder using huggingface_hub

anandjh8
8f310901y ago

Upload folder using huggingface_hub

anandjh8
35d26b71y ago

initial commit

anandjh8