Aniket200325/coder-pretrain-60gb
Coding LLM Pretraining Corpus Built with build_pretrain_dataset.ipynb + finalize_local.py (DuckDB). Composition Source Collected rows Collected size Final rows Final size code 8,415,197 30.00 GB 8,413,850 29.88 GB web 3,492,492 15.00 GB 3,390,372 14.50 GB math 659,720 3.00 GB 658,089 2.95 GB wiki 519,584 2.00 GB 519,535 1.99 GB docs 4,504,069 10.00 GB 4,289,545 9.74 GB Total final corpus: 59.06 GB of raw text (17,271,391 documents) across… See the full description on the dataset page: https://huggingface.co/datasets/Aniket200325/coder-pretrain-60gb.
Coding LLM Pretraining Corpus
Built with build_pretrain_dataset.ipynb + finalize_local.py (DuckDB).
Composition
Total final corpus: 59.06 GB of raw text (17,271,391 documents) across 159 train shards and 1 validation shards. 91,189 documents were removed by dedup/decontamination.
