CoolFace
Datasetpublic

Aniket200325/coder-pretrain-60gb

Coding LLM Pretraining Corpus Built with build_pretrain_dataset.ipynb + finalize_local.py (DuckDB). Composition Source Collected rows Collected size Final rows Final size code 8,415,197 30.00 GB 8,413,850 29.88 GB web 3,492,492 15.00 GB 3,390,372 14.50 GB math 659,720 3.00 GB 658,089 2.95 GB wiki 519,584 2.00 GB 519,535 1.99 GB docs 4,504,069 10.00 GB 4,289,545 9.74 GB Total final corpus: 59.06 GB of raw text (17,271,391 documents) across… See the full description on the dataset page: https://huggingface.co/datasets/Aniket200325/coder-pretrain-60gb.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes133downloads
Dataset Card

Coding LLM Pretraining Corpus

Built with build_pretrain_dataset.ipynb + finalize_local.py (DuckDB).

Composition

SourceCollected rowsCollected sizeFinal rowsFinal size
code8,415,19730.00 GB8,413,85029.88 GB
web3,492,49215.00 GB3,390,37214.50 GB
math659,7203.00 GB658,0892.95 GB
wiki519,5842.00 GB519,5351.99 GB
docs4,504,06910.00 GB4,289,5459.74 GB

Total final corpus: 59.06 GB of raw text (17,271,391 documents) across 159 train shards and 1 validation shards. 91,189 documents were removed by dedup/decontamination.