datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CPT_Data_Pool
CPT Data Pool
This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training.
For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo.
Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.fineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on
cultural knowledge in figurative language, built from the highest-quality
tier of opencsg/Fineweb-Edu-Chinese-V2.1.
Each document is educational Chinese text containing at least one culturally
vetted chengyu, with an appended 【成语注释】 knowledge block listing every
matched idiom's figurative meaning(s) and classical source citation.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.cpt_instruction_datasets
Instruction datasets
Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3.
Dataset creation
Datasets were created using two different techniques:
Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.Theory_SOONibus_1_for_CPT
THEORY SOONibus (#1)
May be used for Continuous Pretraining, such as via this NOTEBOOK
An eclectic selection from the library of books, articles, papers, and varied textual curios I've amassed over the years.
Contains works in English, Russian, and French, including a wealth of critical theory, philosophy, translation theory, comparative literature, literary ethics, radical/revolutionary politics (mainly Marxist-Leninist, Libertarian Communist, Anarchist, Situationist… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Theory_SOONibus_1_for_CPT.pinchbench-clawd
PinchBench Clawd Training Data
Synthetic fine-tuning dataset for training an LLM to act as Clawd, an autonomous AI agent on the OpenClaw framework. Targets the PinchBench benchmark (23 tasks).
Dataset Description
Each example is a multi-turn conversation where Clawd uses tools (file I/O, web search, email, calendar, image generation, memory, etc.) to complete a real-world task. Generated using Claude via the Anthropic Batch API, scored by an LLM judge (1-5), and filtered… See the full description on the dataset page: https://huggingface.co/datasets/cptekur/pinchbench-clawd.Prosperity-Family-Alya-CPT-3B-Kumru
Prosperity Family Alya CPT 3B — Kumru Tokenized
Status: finished - 2B remain - 1B
Developer: Prosperity AIModel/tokenizer: Efe2898/Prosperity-Family-Alya-BaseSource dataset: moganai/turkishfineweb2-cleaned
Tokenizer
Revision: df12a3c9e14c80d1464a8d7a2d624b2b021ab283Vocabulary: 50,176EOS token ID: 3
Filtering
language_score >= 0.98
fasttext_clean_score >= 0.75
document chars: 200 .. 500000
deterministic stream shuffle seed: 20260906
shuffle buffer:… See the full description on the dataset page: https://huggingface.co/datasets/Efe2898/Prosperity-Family-Alya-CPT-3B-Kumru.cpt-coder-sft
CPT / HCPCS Procedure Coder
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Procedure descriptions → correct CPT/HCPCS code with RVU data
Why download this
Build procedure coding assistants, verify CPT code assignments, or automate outpatient charge capture. Covers all specialties in the CMS PFS.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/cpt-coder-sft.Orion-CPT-traindata-v2607serendip-cpt-sinhala
Serendib LLM CPT Sinhala Corpus
A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for
Continual Pre-Training (CPT) of large language models. This dataset was used to adapt
Meta-LLaMA-3-8B to the Sinhala language domain as part of the
Serendib LLM Honours Degree Research Project
at the University of Central Lancashire (UCLan), 2025–2026.
This is one of the largest openly published Sinhala NLP corpora available, containing
23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.mc4-zh-idiom-cpt
mC4 zh — Idiom-Tagged Continued-Pretraining Corpus
A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in
figurative language. Each document is natural web text (from the C4/mC4 zh subset)
containing at least one culturally meaningful chengyu, with an appended knowledge
block that lists every matched idiom together with its figurative meaning(s) and
classical source citation.
Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.cpt-dataset
Hyperswitch CPT Dataset
A comprehensive Continual Pre-Training (CPT) dataset for the Hyperswitch payment processing platform, combining documentation with actual code to build a "world model" understanding of the codebase.
Dataset Description
This dataset was created by mining the Hyperswitch repository and combining it with DeepWiki documentation. It teaches models:
Repository Structure - Where different types of code live
Concept-to-Code Mapping - How abstract concepts… See the full description on the dataset page: https://huggingface.co/datasets/archit11/cpt-dataset.NexaFlow-CPT-Datasetgemma4-31b-cpt-datacode-cpt-corpus
Dataset Card for code-cpt-corpus
Dataset Summary
ဒီ dataset က code-cpt-corpus အတွက် ဖန်တီးထားတာပါ။
Languages
Myanmar (my) / English (en)
Dataset Structure
Data Instances
{ "text": "နမူနာ စာသား", "label": "အညွှန်း" }
Data Fields
text: main content, label: optional.
Data Splits
Split
Files
train
data/train.jsonl
Licensing Information
ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/code-cpt-corpus.L40S-MonEspaceSante-CPT-corpus
Mon Espace Santé — Corpus CPT synthétique
Corpus synthétique de continued pre-training (CPT) généré à partir de 88 faits réels (paires
Q/R « grounded ») de la FAQ du service public français Mon espace santé
(fenyo/MonEspaceSante-FAQ-QA, source == "real").
But : injecter ces connaissances dans les poids d'un LLM sans RAG. Ce corpus sert au CPT décrit dans
le modèle fenyo/L40S-Qwen3-8B-MonEspaceSante-CPT-SFT.
Génération (résumé)
Générateur : Qwen3-32B-FP8 servi par… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/L40S-MonEspaceSante-CPT-corpus.zamai-pashto-clean-cpt
ZamAI Pashto Clean CPT
This dataset is a hyper-cleaned, optimized, and fully deduplicated version of tasal9/ZamAI-Pashto-Mega-Dataset intended for Causal Language Modeling (CLM), pre-training, or fine-tuning text models in the Pashto language.
🛠️ Pipeline & Filtering Details
Before processing the ~1.5 GB stream, an Internal Built-In Self-Test (BIST) was executed to verify environment I/O permissions, validate character ratio logic, and check connection… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/zamai-pashto-clean-cpt.indonesian-dfk-cpt-dataset
Indonesian DFK Domain CPT Corpus
Dataset Description
Dataset ini merupakan korpus teks Bahasa Indonesia untuk kebutuhan Continued Pre-Training atau CPT pada domain DFK, yaitu domain yang berkaitan dengan topik-topik yang sering menjadi sasaran disinformasi, fitnah, dan kebencian di Indonesia.
Istilah DFK dalam dataset ini tidak berarti bahwa teks berisi disinformasi, fitnah, atau ujaran kebencian. DFK di sini merujuk pada domain atau topik yang sering menjadi sasaran DFK… See the full description on the dataset page: https://huggingface.co/datasets/aikathata/indonesian-dfk-cpt-dataset.beckett-dramatic-works-cpt
Samuel Beckett Complete Dramatic Works (CPT Dataset)
This dataset contains the complete, raw theatrical texts of Samuel Beckett's dramatic works.
It was specifically compiled for Continued Pre-Training (CPT) to teach Large Language Models the distinct vocabulary, pacing, and minimalist stage directions characteristic of Beckett's writing style.
Contents
This dataset consists of raw text blocks extracted from:
Waiting for Godot
Endgame
Krapp's Last Tape
Happy… See the full description on the dataset page: https://huggingface.co/datasets/alst10/beckett-dramatic-works-cpt.rosetta-ko-law-synth-cpt
rosetta-ko-law-synth-cpt
Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Law Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-law-synth-cpt (this repo)
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-cpt.rosetta-ko-tourism-synth-cpt
rosetta-ko-tourism-synth-cpt
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt (this repo)
continued-pretraining corpus (plain… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-cpt.nyxmed-icd-cpt-8kAsylum-Final-CPT-AdapterLFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627
LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627
Full Korean CPT mix converted to LFM-style text JSONL, about 4B-token training source.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627.legal-chunks-cpt
Legal Document Chunks for Continued Pretraining
This dataset contains 1329 legal document chunks extracted from various legal documents across multiple jurisdictions. Each chunk is enriched with comprehensive metadata labels for filtering, analysis, and domain-specific training.
Dataset Information
Total Chunks: 1,329
Format: jsonl-text
Sorted: By document ID and chunk index (maintains document continuity)
Source: Legal documents processed through enhanced parser with… See the full description on the dataset page: https://huggingface.co/datasets/rzeraat/legal-chunks-cpt.Genshin-CPT-Public
License & Citation
This dataset is a derivative work based on mrzjy/multimodal-genshin-impact.
Original License: CC-BY-SA (Creative Commons Attribution-ShareAlike)
Citation / Attribution:
This processed dataset is derived from the work of mrzjy. As per the CC-BY-SA license, this derivative work is also released under the same license.
Original Dataset Citation:
mrzjy. (2023). multimodal-genshin-impact [Dataset]. Hugging Face.… See the full description on the dataset page: https://huggingface.co/datasets/NaruseShiroha/Genshin-CPT-Public.asm-cpt-mixedTL-001-CPT-Final-RTTHyperSwitch-Repo-CPT-Dataset-v2
Hyperswitch Rust Codebase Dataset
A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models.
📊 Dataset Overview
This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset-v2.H100-MonEspaceSante-CPT-corpus
🔧 Code & reproduction complète (scripts, RUNBOOK, reproduce.sh, conversations) : https://github.com/AlexandreFenyo/MonEspaceSante-H100-reproduction
Mon espace santé — Corpus synthétique pour CPT (EntiGraph, FR)
Corpus synthétique en français pour le continued pre-training (CPT) d'un LLM, destiné à
injecter dans les poids les connaissances de la FAQ du service public « Mon espace santé ».
Provenance (1-hop, anti model-collapse)
Généré uniquement à partir des 88… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/H100-MonEspaceSante-CPT-corpus.HyperSwitch-Repo-CPT-Dataset
Hyperswitch Rust Codebase Dataset
A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models.
📊 Dataset Overview
This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset.
