CoolFace
Datasetpublicgated

Gugu8/Scraped-Data

Scraped-Data — Building to 1 TB A continuously-growing, cleaned text corpus scraped, parsed, cleaned, losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline. Current repo size: ~26 GB → target: 1 TB (incremental batch uploads). Current files File Size Contents corpus_batch1.clean.zip 4.80 GB ~2.9M Wikipedia docs (stored, from first scrape) corpus_batch2.clean.zip 4.85 GB ~2.5M docs corpus_batch3.clean.zip 4.91 GB ~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.

sourceHugging Faceotherupdated 18d agoView on Hugging Face
0likes1.8kdownloads
settings

This repository belongs to Gugu8 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameScraped-Data
visibilitypublic
licenceother
gatedyes
ownerGugu8
Account settings
Gugu8/Scraped-Data · CoolFace