datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CPT_Data_Pool
CPT Data Pool
This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training.
For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo.
Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.fineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on
cultural knowledge in figurative language, built from the highest-quality
tier of opencsg/Fineweb-Edu-Chinese-V2.1.
Each document is educational Chinese text containing at least one culturally
vetted chengyu, with an appended 【成语注释】 knowledge block listing every
matched idiom's figurative meaning(s) and classical source citation.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.dim58-cpuData-31cases
Dim58 CPU Data — 31 Cases
Dataset uploaded from:
/mnt/data/ubuntu/research/outputs/data_cpu_geodesic58
Dataset summary
Property
Value
Repository
hosseinbv/dim58-cpuData-31cases
Number of files
64
Total size
17.87 GB
Source folder
data_cpu_geodesic58
File types
Extension
File count
.npz
62
.json
1
.csv
1
Top-level contents
0000_internal_case1_data.npz
0001_internal_B_10.npz… See the full description on the dataset page: https://huggingface.co/datasets/hosseinbv/dim58-cpuData-31cases.spec_cpu_branch_tracesCpp-Code-LargeCpp-Code-Large
Cpp-Code-Large is a large-scale corpus of C++ source code comprising more than 5 million lines of C++ code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the C++ ecosystem.
By providing a high-volume, language-specific corpus, Cpp-Code-Large enables systematic experimentation in C++-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Cpp-Code-Large.step35-en2pl-conv-pass4-jsonlconversations: 1,251,034
chat-template tokens (role+content, incl. special tokens): 2,664,206,408
reasoning_content tokens (not covered by chat template, counted separately): 6,662,763,429
avg tokens/conversation: 2129.6
used tokenizer: APT4
CPRet-Embeddings
CPRet-Embeddings
This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server.
You can explore the retrieval server via the online demo at https://cpret.online/.
📦 Files
probs_2609.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-Embeddings.step35-en2pl-conv-pass5-jsonlLiveCodeBench-CPP
LiveCodeBench-CPP: An Extension of LiveCodeBench for Contamination Free Evaluation in C++
Overview
LiveCodeBench-CPP includes 454 problems from the release_v6 of LiveCodeBench, covering the period from October 2024 to May 2025. These problems are sourced from AtCoder (287 problems) and LeetCode (167 problems).
AtCoder Problems: These require generated solutions to read inputs from standard input (stdin) and write outputs to standard output (stdout). For unit testing, the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LiveCodeBench-CPP.step35-en2pl-conv-pass2-jsonlcpt_instruction_datasets
Instruction datasets
Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3.
Dataset creation
Datasets were created using two different techniques:
Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.step35-en2pl-conv-pass7-jsonliran_inflation_and_cpi_1936_2022
⚠️ نسخهٔ جایگزین
این مجموعهداده با روششناسیِ فعلیِ فرمانا بهروز نمیشود.
→ Farmaanaa/iran_cpi_and_inflation_multisource
فایلهای قبلی برای آرشیو در دسترس میمانند.
— farmaanaa.ir
CPsyCoun
CPsyCounD
The high-quality multi-turn dialogue dataset, which has a total of 3,134 multi-turn consultation dialogues. CPsyCounD covers nine representative topics and seven classic schools of psychological counseling.
Paper: CPsyCoun
Data analysis
Topic types
Self-growth
Emotion&Stress
Education
Love&Marriage
Family Relationship
Social Relationship
Sex
Career
Mental Disease
Consulting schools
Psychoanalytic Therapy
Cognitive Behavioral Therapy… See the full description on the dataset page: https://huggingface.co/datasets/CAS-SIAT-XinHai/CPsyCoun.CPRet-Embeddings
CPRet-Embeddings
This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server.
You can explore the retrieval server via the online demo at https://cpret.online/.
📦 Files
probs_2606.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/upctanker/CPRet-Embeddings.magenta-realtime-mlx-cpp
Magenta RealTime — C++ MLX runtime bundle
This dataset is a re-packaging of
Google's Magenta RealTime weights
for the C++ MLX runtime in
rhymeswithlion/magenta-realtime-mlx-cpp.
It contains exactly what mlx-stream needs at startup; nothing more, nothing
less. The upstream .pt / .npy checkpoints are intentionally not
mirrored here — they're only useful for the (Python) re-export tooling on the
project's main distribution.
Contents
.
├──… See the full description on the dataset page: https://huggingface.co/datasets/rhymeswithlion/magenta-realtime-mlx-cpp.glm-moe-dsa-tiny-cpu-repro-v1
Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay
Reproducibility evidence for
malaiwah/glm-moe-dsa-tiny-random-bf16,
checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37.
This is a synthetic pipeline test, not a quality benchmark, quantization measurement,
qualified production reference, or registry submission. The model is random-init.
No GPU or paid cloud job was used.
Observed result
Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.CP-Bench
[!IMPORTANT]
CP-Bench has been superseded by DCP-Bench-Open. This repository is kept for archival and reproducibility purposes only.
Please use the current version here: DCP-Bench/DCP-Bench-Open.
CP-Bench: A dataset for evaluating LLM-driven constraint modelling
This dataset is designed to facilitate the evaluation of LLM-based methods for translating natural language problem descriptions into accurate constraint specifications. It contains diverse combinatorial problems, and is… See the full description on the dataset page: https://huggingface.co/datasets/kostis-init/CP-Bench.Theory_SOONibus_1_for_CPT
THEORY SOONibus (#1)
May be used for Continuous Pretraining, such as via this NOTEBOOK
An eclectic selection from the library of books, articles, papers, and varied textual curios I've amassed over the years.
Contains works in English, Russian, and French, including a wealth of critical theory, philosophy, translation theory, comparative literature, literary ethics, radical/revolutionary politics (mainly Marxist-Leninist, Libertarian Communist, Anarchist, Situationist… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Theory_SOONibus_1_for_CPT.cpp-code-code_search_net-style
C++ Dataset
documentation source: https://huggingface.co/docs/datasets/main/en/repository_structure
Supported Tasks and Leaderboards
language-modeling: The dataset can be used to train a model for modelling programming languages, which consists in building language models for programming languages.
Language
C++ programming language
Dataset Structure
Data Instances
A data point consists of a function code along with its documentation.… See the full description on the dataset page: https://huggingface.co/datasets/malteklaes/cpp-code-code_search_net-style.pinchbench-clawd
PinchBench Clawd Training Data
Synthetic fine-tuning dataset for training an LLM to act as Clawd, an autonomous AI agent on the OpenClaw framework. Targets the PinchBench benchmark (23 tasks).
Dataset Description
Each example is a multi-turn conversation where Clawd uses tools (file I/O, web search, email, calendar, image generation, memory, etc.) to complete a real-world task. Generated using Claude via the Anthropic Batch API, scored by an LLM judge (1-5), and filtered… See the full description on the dataset page: https://huggingface.co/datasets/cptekur/pinchbench-clawd.Prosperity-Family-Alya-CPT-3B-Kumru
Prosperity Family Alya CPT 3B — Kumru Tokenized
Status: finished - 2B remain - 1B
Developer: Prosperity AIModel/tokenizer: Efe2898/Prosperity-Family-Alya-BaseSource dataset: moganai/turkishfineweb2-cleaned
Tokenizer
Revision: df12a3c9e14c80d1464a8d7a2d624b2b021ab283Vocabulary: 50,176EOS token ID: 3
Filtering
language_score >= 0.98
fasttext_clean_score >= 0.75
document chars: 200 .. 500000
deterministic stream shuffle seed: 20260906
shuffle buffer:… See the full description on the dataset page: https://huggingface.co/datasets/Efe2898/Prosperity-Family-Alya-CPT-3B-Kumru.cpos
CPOS-HG Training Corpora
Training and validation corpora used in the cross-lingual poverty-of-stimulus
(CPOS) experiments reported in Once a Tree, Always a Tree? Cross-lingual
Transfer of Hierarchical Generalization in Language Models.
Configurations
The L1 configurations cross four languages with two evidence conditions:
*_l1_ambiguous: hierarchical-evidence target ratio 0.000.
*_l1_disambiguating: hierarchical-evidence target ratio 0.500.
English L2 is fixed… See the full description on the dataset page: https://huggingface.co/datasets/shin0729/cpos.t_cppswe-agent-cpu-dynamorio-pilot-sympy-15599
One-agent CPU trace pilot
Preliminary research data. Validation is incomplete; this is not a confirmed dead-state result.
One live mini-SWE-agent 2.4.6 execution of sympy__sympy-15599, using a separate Qwen3-Coder-30B-A3B-Instruct AWQ server. The collector finished normally in 907.94 seconds. The agent made 57 model calls and submitted a patch; benchmark evaluation was not run. This is mini-SWE-agent, not the original full SWE-agent implementation.
What is included… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/swe-agent-cpu-dynamorio-pilot-sympy-15599.cpt-coder-sft
CPT / HCPCS Procedure Coder
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Procedure descriptions → correct CPT/HCPCS code with RVU data
Why download this
Build procedure coding assistants, verify CPT code assignments, or automate outpatient charge capture. Covers all specialties in the CMS PFS.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/cpt-coder-sft.poziomka-fun-v9-jsonlmc4-zh-idiom-cpt
mC4 zh — Idiom-Tagged Continued-Pretraining Corpus
A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in
figurative language. Each document is natural web text (from the C4/mC4 zh subset)
containing at least one culturally meaningful chengyu, with an appended knowledge
block that lists every matched idiom together with its figurative meaning(s) and
classical source citation.
Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.Orion-CPT-traindata-v2607serendip-cpt-sinhala
Serendib LLM CPT Sinhala Corpus
A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for
Continual Pre-Training (CPT) of large language models. This dataset was used to adapt
Meta-LLaMA-3-8B to the Sinhala language domain as part of the
Serendib LLM Honours Degree Research Project
at the University of Central Lancashire (UCLan), 2025–2026.
This is one of the largest openly published Sinhala NLP corpora available, containing
23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.
