datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on
cultural knowledge in figurative language, built from the highest-quality
tier of opencsg/Fineweb-Edu-Chinese-V2.1.
Each document is educational Chinese text containing at least one culturally
vetted chengyu, with an appended 【成语注释】 knowledge block listing every
matched idiom's figurative meaning(s) and classical source citation.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.cpt_instruction_datasets
Instruction datasets
Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3.
Dataset creation
Datasets were created using two different techniques:
Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.pinchbench-clawd
PinchBench Clawd Training Data
Synthetic fine-tuning dataset for training an LLM to act as Clawd, an autonomous AI agent on the OpenClaw framework. Targets the PinchBench benchmark (23 tasks).
Dataset Description
Each example is a multi-turn conversation where Clawd uses tools (file I/O, web search, email, calendar, image generation, memory, etc.) to complete a real-world task. Generated using Claude via the Anthropic Batch API, scored by an LLM judge (1-5), and filtered… See the full description on the dataset page: https://huggingface.co/datasets/cptekur/pinchbench-clawd.cpt-coder-sft
CPT / HCPCS Procedure Coder
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Procedure descriptions → correct CPT/HCPCS code with RVU data
Why download this
Build procedure coding assistants, verify CPT code assignments, or automate outpatient charge capture. Covers all specialties in the CMS PFS.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/cpt-coder-sft.serendip-cpt-sinhala
Serendib LLM CPT Sinhala Corpus
A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for
Continual Pre-Training (CPT) of large language models. This dataset was used to adapt
Meta-LLaMA-3-8B to the Sinhala language domain as part of the
Serendib LLM Honours Degree Research Project
at the University of Central Lancashire (UCLan), 2025–2026.
This is one of the largest openly published Sinhala NLP corpora available, containing
23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.mc4-zh-idiom-cpt
mC4 zh — Idiom-Tagged Continued-Pretraining Corpus
A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in
figurative language. Each document is natural web text (from the C4/mC4 zh subset)
containing at least one culturally meaningful chengyu, with an appended knowledge
block that lists every matched idiom together with its figurative meaning(s) and
classical source citation.
Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.cpt-dataset
Hyperswitch CPT Dataset
A comprehensive Continual Pre-Training (CPT) dataset for the Hyperswitch payment processing platform, combining documentation with actual code to build a "world model" understanding of the codebase.
Dataset Description
This dataset was created by mining the Hyperswitch repository and combining it with DeepWiki documentation. It teaches models:
Repository Structure - Where different types of code live
Concept-to-Code Mapping - How abstract concepts… See the full description on the dataset page: https://huggingface.co/datasets/archit11/cpt-dataset.L40S-MonEspaceSante-CPT-corpus
Mon Espace Santé — Corpus CPT synthétique
Corpus synthétique de continued pre-training (CPT) généré à partir de 88 faits réels (paires
Q/R « grounded ») de la FAQ du service public français Mon espace santé
(fenyo/MonEspaceSante-FAQ-QA, source == "real").
But : injecter ces connaissances dans les poids d'un LLM sans RAG. Ce corpus sert au CPT décrit dans
le modèle fenyo/L40S-Qwen3-8B-MonEspaceSante-CPT-SFT.
Génération (résumé)
Générateur : Qwen3-32B-FP8 servi par… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/L40S-MonEspaceSante-CPT-corpus.indonesian-dfk-cpt-dataset
Indonesian DFK Domain CPT Corpus
Dataset Description
Dataset ini merupakan korpus teks Bahasa Indonesia untuk kebutuhan Continued Pre-Training atau CPT pada domain DFK, yaitu domain yang berkaitan dengan topik-topik yang sering menjadi sasaran disinformasi, fitnah, dan kebencian di Indonesia.
Istilah DFK dalam dataset ini tidak berarti bahwa teks berisi disinformasi, fitnah, atau ujaran kebencian. DFK di sini merujuk pada domain atau topik yang sering menjadi sasaran DFK… See the full description on the dataset page: https://huggingface.co/datasets/aikathata/indonesian-dfk-cpt-dataset.rosetta-ko-law-synth-cpt
rosetta-ko-law-synth-cpt
Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Law Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-law-synth-cpt (this repo)
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-cpt.rosetta-ko-tourism-synth-cpt
rosetta-ko-tourism-synth-cpt
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt (this repo)
continued-pretraining corpus (plain… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-cpt.legal-chunks-cpt
Legal Document Chunks for Continued Pretraining
This dataset contains 1329 legal document chunks extracted from various legal documents across multiple jurisdictions. Each chunk is enriched with comprehensive metadata labels for filtering, analysis, and domain-specific training.
Dataset Information
Total Chunks: 1,329
Format: jsonl-text
Sorted: By document ID and chunk index (maintains document continuity)
Source: Legal documents processed through enhanced parser with… See the full description on the dataset page: https://huggingface.co/datasets/rzeraat/legal-chunks-cpt.H100-MonEspaceSante-CPT-corpus
🔧 Code & reproduction complète (scripts, RUNBOOK, reproduce.sh, conversations) : https://github.com/AlexandreFenyo/MonEspaceSante-H100-reproduction
Mon espace santé — Corpus synthétique pour CPT (EntiGraph, FR)
Corpus synthétique en français pour le continued pre-training (CPT) d'un LLM, destiné à
injecter dans les poids les connaissances de la FAQ du service public « Mon espace santé ».
Provenance (1-hop, anti model-collapse)
Généré uniquement à partir des 88… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/H100-MonEspaceSante-CPT-corpus.HyperSwitch-Repo-CPT-Dataset
Hyperswitch Rust Codebase Dataset
A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models.
📊 Dataset Overview
This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset.legal-chunks-cpt-test
Legal Document Chunks for Continued Pretraining
This dataset contains 1281 legal document chunks extracted from various legal documents across multiple jurisdictions. Each chunk is enriched with comprehensive metadata labels for filtering, analysis, and domain-specific training.
Dataset Information
Total Chunks: 1,281
Format: jsonl-text
Sorted: By document ID and chunk index (maintains document continuity)
Source: Legal documents processed through enhanced parser with… See the full description on the dataset page: https://huggingface.co/datasets/rzeraat/legal-chunks-cpt-test.HyperSwitch-Repo-CPT-Dataset-v2
Hyperswitch Rust Codebase Dataset
A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models.
📊 Dataset Overview
This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset-v2.German-RAG-CPT-HESSIAN-AI
German-RAG-CPT (Continued Pre-Training) Tasks Dataset
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The CPT Tasks Dataset is a comprehensive collection designed for continued pre-training of language models, focusing on three core competencies: context-based question answering, structured reasoning, and summarization. The dataset comprises approximately 620,000 examples, with 420,000 in German and 200,000 in English.
Developed by Avemio AG… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-CPT-HESSIAN-AI.Continued-Pre-Training-CPT-Paite
Continued-Pre-Training-CPT-Paite (Master Collection)
This repository contains the unified, high-density raw text data used for the Continued Pre-Training (CPT) of the Sensix Paite models (Gemma-4-31B Master and Gemma-4-2B/5B Nitro).
The dataset is specifically designed to expand a base model's vocabulary and internalize Paite linguistic patterns, syntax, and tonal logic before moving to instruction fine-tuning (SFT).
Dataset Composition
This is a unified dataset… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-CPT-Paite.gold-silver-mineral-process-cpt-candidates
Gold/silver mineral-process CPT candidates
English raw documents (text) about gold/silver and transferable hard-rock mineral processing.
Source: BAAI/IndustryCorpus2_mining revision bf358a2f8105e4ac468141796e5a1a530685ae2e. English only.
These are documents, not chat pairs.
Configs
Config
Rows
Notes
default
59,749
all English bands
english_high
18,864
publisher quality 4.00–4.59
english_middle
34,601
publisher quality 3.00–4.00
english_low
6,284… See the full description on the dataset page: https://huggingface.co/datasets/hicham-taoufik/gold-silver-mineral-process-cpt-candidates.jwiki_cpt
Dataset Card for J Wiki Continuous Pretraining
Markdown documents converted from the Jsoftware wiki for continuous pretraining (next-token prediction) of models that should read and write the J programming language.
Each row is one wiki page. J session logs and scripts are fenced as j code blocks. Leading spaces, _ negatives (for example 3j_5, _1), and boxed-array characters are kept.
Dataset Details
Dataset Sources
Wiki:… See the full description on the dataset page: https://huggingface.co/datasets/otisberg/jwiki_cpt.
