kaptaan45/KapCode-1B
KapCode-1B: Curated 1-Billion Token Dataset for Compact Code Models KapCode-1B is a high-quality, 1-billion-token curated dataset designed for Continued Pre-Training (CPT) and domain adaptation of compact Large Language Models. Engineered specifically to empower models under 1 billion parameters with robust code generation, technical comprehension, mathematical reasoning, and Fill-in-the-Middle (FIM) infilling capabilities, KapCode-1B combines multi-lingual code… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapCode-1B.
029
This repository is gated, so its file contents are only served once you have accepted the publisher's terms at Hugging Face. Open it at the source above.
