CoolFace
Datasetpublicgated

kaptaan45/KapCode-1B

KapCode-1B: Curated 1-Billion Token Dataset for Compact Code Models KapCode-1B is a high-quality, 1-billion-token curated dataset designed for Continued Pre-Training (CPT) and domain adaptation of compact Large Language Models. Engineered specifically to empower models under 1 billion parameters with robust code generation, technical comprehension, mathematical reasoning, and Fill-in-the-Middle (FIM) infilling capabilities, KapCode-1B combines multi-lingual code… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapCode-1B.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes29downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.