CoolFace
Datasetpublicgated

kaptaan45/KapCode-1B

KapCode-1B: Curated 1-Billion Token Dataset for Compact Code Models KapCode-1B is a high-quality, 1-billion-token curated dataset designed for Continued Pre-Training (CPT) and domain adaptation of compact Large Language Models. Engineered specifically to empower models under 1 billion parameters with robust code generation, technical comprehension, mathematical reasoning, and Fill-in-the-Middle (FIM) infilling capabilities, KapCode-1B combines multi-lingual code… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapCode-1B.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes29downloads
.gitattributesDownload Raw Back to root

This repository is gated, so its file contents are only served once you have accepted the publisher's terms at Hugging Face. Open it at the source above.