ZhejiangLab/CPT_Data_Pool
CPT Data Pool This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training. For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo. Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.
Update README.md
Upload LICENSE.md
Create README.md
Delete pes2o
Delete arxiv
Upload folder using huggingface_hub
Delete code/readme.md
Delete reddit
Delete megawika
Delete flan
Delete cc_news
Delete books
Delete wiki
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
chore: relocate math sub-folders
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Delete code/readme.md
chore: move code files to subdirectory
强制迁移第 7 组数据
强制迁移第 6 组数据
强制迁移第 5 组数据
强制迁移第 4 组数据
强制迁移第 3 组数据
强制迁移第 2 组数据
强制迁移第 1 组数据
强制归位:将第 7 组文件移入 code/
强制归位:将第 6 组文件移入 code/
强制归位:将第 5 组文件移入 code/
强制归位:将第 4 组文件移入 code/
强制归位:将第 3 组文件移入 code/
强制归位:将第 2 组文件移入 code/
强制归位:将第 1 组文件移入 code/
整理数据:批量移动第 4 组文件 (共 20 个)
整理数据:批量移动第 3 组文件 (共 100 个)
整理数据:批量移动第 2 组文件 (共 100 个)
整理数据:批量移动第 1 组文件 (共 100 个)
Create code/readme.md
Add files using upload-large-folder tool
Add files using upload-large-folder tool
