CoolFace
Datasetpublic

ZhejiangLab/CPT_Data_Pool

CPT Data Pool This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training. For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo. Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
1likes4.7kdownloads
settings

This repository belongs to ZhejiangLab on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameCPT_Data_Pool
visibilitypublic
licenceapache-2.0
gatedno
ownerZhejiangLab
Account settings
ZhejiangLab/CPT_Data_Pool · CoolFace