datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cpt_instruction_datasets
Instruction datasets
Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3.
Dataset creation
Datasets were created using two different techniques:
Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.wiki-events-cpt
Wikipedia Events CPT
jhdlee/wiki-events-cpt is a public research dataset of 150 selected English Wikipedia articles with compact metadata for continual pretraining (CPT).
Split
Articles
Event window (end exclusive)
cohort_a
75
2023-01-01 to 2024-10-01
cohort_b
75
2024-10-01 to 2025-09-01
Each cohort has 25 articles per topic: natural_hazards, elections, and sports. Cohorts group events by their reviewed whole-occurrence intervals; they are not… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/wiki-events-cpt.cpt-dataset
Hyperswitch CPT Dataset
A comprehensive Continual Pre-Training (CPT) dataset for the Hyperswitch payment processing platform, combining documentation with actual code to build a "world model" understanding of the codebase.
Dataset Description
This dataset was created by mining the Hyperswitch repository and combining it with DeepWiki documentation. It teaches models:
Repository Structure - Where different types of code live
Concept-to-Code Mapping - How abstract concepts… See the full description on the dataset page: https://huggingface.co/datasets/archit11/cpt-dataset.MSA-cpt-100b
MSA CPT corpus (cpt_100b)
Continual-pre-training corpus for Memory Sparse Attention (MSA) / Generative
Retrieval on a Qwen3.5 backbone. Built 2026-06-29 by aggregating and
within-source-deduplicating 45 public retrieval / QA datasets into a unified
query↔document contract.
Layout
corpus/ # document side (.jsonl.gz shards)
train/ # query↔positive-doc training pairs (.jsonl.gz shards)
manifest.json
Stats (from manifest.json)
field
value… See the full description on the dataset page: https://huggingface.co/datasets/william-0g/MSA-cpt-100b.German-RAG-CPT-HESSIAN-AI
German-RAG-CPT (Continued Pre-Training) Tasks Dataset
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The CPT Tasks Dataset is a comprehensive collection designed for continued pre-training of language models, focusing on three core competencies: context-based question answering, structured reasoning, and summarization. The dataset comprises approximately 620,000 examples, with 420,000 in German and 200,000 in English.
Developed by Avemio AG… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-CPT-HESSIAN-AI.
