anandjh8/common-crawl-english-filtered
🧠 FineWeb-English-Filtered 📘 Dataset Summary FineWeb-English-Filtered is a large-scale, cleaned, English-only text dataset derived from Common Crawl’s WET archives.It contains 940 million documents of publicly available web text, converted into Apache Parquet format with a consistent schema for fast and efficient data loading. The dataset was generated using a custom AWS Glue pipeline that processed, filtered, and merged .wet files across multiple terabytes of… See the full description on the dataset page: https://huggingface.co/datasets/anandjh8/common-crawl-english-filtered.
2324
Update README.md
Create README.md
Upload folder using huggingface_hub
Upload folder using huggingface_hub
initial commit
