CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZhejiangLab /CPT_Data_Pool CPT Data Pool This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training. For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo. Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.text100M<n<1B0 likes4.9k downloads3mo agoHugging Face02jiviteshjn /fineweb-edu-zh-chengyu-cpt Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.tabulartext-generation1M<n<10M1 likes1.9k downloads2mo agoHugging Face03proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes338 downloads5mo agoHugging Face04AlekseyCalvin /Theory_SOONibus_1_for_CPT THEORY SOONibus (#1) May be used for Continuous Pretraining, such as via this NOTEBOOK An eclectic selection from the library of books, articles, papers, and varied textual curios I've amassed over the years. Contains works in English, Russian, and French, including a wealth of critical theory, philosophy, translation theory, comparative literature, literary ethics, radical/revolutionary politics (mainly Marxist-Leninist, Libertarian Communist, Anarchist, Situationist… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Theory_SOONibus_1_for_CPT.text10K<n<100K0 likes158 downloads7mo agoHugging Face05cptekur /pinchbench-clawd PinchBench Clawd Training Data Synthetic fine-tuning dataset for training an LLM to act as Clawd, an autonomous AI agent on the OpenClaw framework. Targets the PinchBench benchmark (23 tasks). Dataset Description Each example is a multi-turn conversation where Clawd uses tools (file I/O, web search, email, calendar, image generation, memory, etc.) to complete a real-world task. Generated using Claude via the Anthropic Batch API, scored by an LLM judge (1-5), and filtered… See the full description on the dataset page: https://huggingface.co/datasets/cptekur/pinchbench-clawd.texttext-generation1K<n<10K0 likes134 downloads6mo agoHugging Face06Efe2898 /Prosperity-Family-Alya-CPT-3B-Kumru Prosperity Family Alya CPT 3B — Kumru Tokenized Status: finished - 2B remain - 1B Developer: Prosperity AIModel/tokenizer: Efe2898/Prosperity-Family-Alya-BaseSource dataset: moganai/turkishfineweb2-cleaned Tokenizer Revision: df12a3c9e14c80d1464a8d7a2d624b2b021ab283Vocabulary: 50,176EOS token ID: 3 Filtering language_score >= 0.98 fasttext_clean_score >= 0.75 document chars: 200 .. 500000 deterministic stream shuffle seed: 20260906 shuffle buffer:… See the full description on the dataset page: https://huggingface.co/datasets/Efe2898/Prosperity-Family-Alya-CPT-3B-Kumru.tabularn<1K0 likes129 downloads17d agoHugging Face07AmareshHebbar /cpt-coder-sft CPT / HCPCS Procedure Coder Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Procedure descriptions → correct CPT/HCPCS code with RVU data Why download this Build procedure coding assistants, verify CPT code assignments, or automate outpatient charge capture. Covers all specialties in the CMS PFS. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/cpt-coder-sft.texttext-generation10K<n<100K0 likes98 downloads3mo agoHugging Face083tic /Orion-CPT-traindata-v2607text100M<n<1B0 likes79 downloads3mo agoHugging Face09Chamaka8 /serendip-cpt-sinhala Serendib LLM CPT Sinhala Corpus A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for Continual Pre-Training (CPT) of large language models. This dataset was used to adapt Meta-LLaMA-3-8B to the Sinhala language domain as part of the Serendib LLM Honours Degree Research Project at the University of Central Lancashire (UCLan), 2025–2026. This is one of the largest openly published Sinhala NLP corpora available, containing 23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.texttext-generation10M<n<100M0 likes71 downloads6mo agoHugging Face10jiviteshjn /mc4-zh-idiom-cpt mC4 zh — Idiom-Tagged Continued-Pretraining Corpus A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in figurative language. Each document is natural web text (from the C4/mC4 zh subset) containing at least one culturally meaningful chengyu, with an appended knowledge block that lists every matched idiom together with its figurative meaning(s) and classical source citation. Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.tabulartext-generation1M<n<10M0 likes69 downloads2mo agoHugging Face11archit11 /cpt-dataset Hyperswitch CPT Dataset A comprehensive Continual Pre-Training (CPT) dataset for the Hyperswitch payment processing platform, combining documentation with actual code to build a "world model" understanding of the codebase. Dataset Description This dataset was created by mining the Hyperswitch repository and combining it with DeepWiki documentation. It teaches models: Repository Structure - Where different types of code live Concept-to-Code Mapping - How abstract concepts… See the full description on the dataset page: https://huggingface.co/datasets/archit11/cpt-dataset.texttext-generationn<1K0 likes59 downloads11mo agoHugging Face12Venky0705 /NexaFlow-CPT-Datasettextn<1K0 likes55 downloads17d agoHugging Face13AiForgeMaster /gemma4-31b-cpt-datatext100K<n<1M0 likes30 downloads5mo agoHugging Face14kkomyoeminaung /code-cpt-corpus Dataset Card for code-cpt-corpus Dataset Summary ဒီ dataset က code-cpt-corpus အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train data/train.jsonl Licensing Information ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/code-cpt-corpus.text100K<n<1M0 likes30 downloads2mo agoHugging Face15fenyo /L40S-MonEspaceSante-CPT-corpus Mon Espace Santé — Corpus CPT synthétique Corpus synthétique de continued pre-training (CPT) généré à partir de 88 faits réels (paires Q/R « grounded ») de la FAQ du service public français Mon espace santé (fenyo/MonEspaceSante-FAQ-QA, source == "real"). But : injecter ces connaissances dans les poids d'un LLM sans RAG. Ce corpus sert au CPT décrit dans le modèle fenyo/L40S-Qwen3-8B-MonEspaceSante-CPT-SFT. Génération (résumé) Générateur : Qwen3-32B-FP8 servi par… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/L40S-MonEspaceSante-CPT-corpus.texttext-generation10K<n<100K0 likes28 downloads4mo agoHugging Face16nassimjp /zamai-pashto-clean-cpt ZamAI Pashto Clean CPT This dataset is a hyper-cleaned, optimized, and fully deduplicated version of tasal9/ZamAI-Pashto-Mega-Dataset intended for Causal Language Modeling (CLM), pre-training, or fine-tuning text models in the Pashto language. 🛠️ Pipeline & Filtering Details Before processing the ~1.5 GB stream, an Internal Built-In Self-Test (BIST) was executed to verify environment I/O permissions, validate character ratio logic, and check connection… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/zamai-pashto-clean-cpt.text1M<n<10M0 likes28 downloads3mo agoHugging Face17aikathata /indonesian-dfk-cpt-dataset Indonesian DFK Domain CPT Corpus Dataset Description Dataset ini merupakan korpus teks Bahasa Indonesia untuk kebutuhan Continued Pre-Training atau CPT pada domain DFK, yaitu domain yang berkaitan dengan topik-topik yang sering menjadi sasaran disinformasi, fitnah, dan kebencian di Indonesia. Istilah DFK dalam dataset ini tidak berarti bahwa teks berisi disinformasi, fitnah, atau ujaran kebencian. DFK di sini merujuk pada domain atau topik yang sering menjadi sasaran DFK… See the full description on the dataset page: https://huggingface.co/datasets/aikathata/indonesian-dfk-cpt-dataset.texttext-generation1M<n<10M1 likes27 downloads4mo agoHugging Face18alst10 /beckett-dramatic-works-cpt Samuel Beckett Complete Dramatic Works (CPT Dataset) This dataset contains the complete, raw theatrical texts of Samuel Beckett's dramatic works. It was specifically compiled for Continued Pre-Training (CPT) to teach Large Language Models the distinct vocabulary, pacing, and minimalist stage directions characteristic of Beckett's writing style. Contents This dataset consists of raw text blocks extracted from: Waiting for Godot Endgame Krapp's Last Tape Happy… See the full description on the dataset page: https://huggingface.co/datasets/alst10/beckett-dramatic-works-cpt.text1K<n<10K0 likes27 downloads1mo agoHugging Face19PoSTMEDIA /rosetta-ko-law-synth-cptgated rosetta-ko-law-synth-cpt Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Law Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-law-synth-cpt (this repo) continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-cpt.texttext-generation100K<n<1M0 likes22 downloads13d agoHugging Face20PoSTMEDIA /rosetta-ko-tourism-synth-cptgated rosetta-ko-tourism-synth-cpt Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Tourism Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-tourism-synth-cpt (this repo) continued-pretraining corpus (plain… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-cpt.texttext-generation10K<n<100K0 likes22 downloads13d agoHugging Face21vineetdaniels /nyxmed-icd-cpt-8ktext1K<n<10K0 likes20 downloads9mo agoHugging Face22delta34 /Asylum-Final-CPT-Adaptertextn<1K1 likes20 downloads3mo agoHugging Face23LLM-OS-Models /LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627 LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627 Full Korean CPT mix converted to LFM-style text JSONL, about 4B-token training source. This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow. Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627.text1M<n<10M0 likes18 downloads3mo agoHugging Face24rzeraat /legal-chunks-cpt Legal Document Chunks for Continued Pretraining This dataset contains 1329 legal document chunks extracted from various legal documents across multiple jurisdictions. Each chunk is enriched with comprehensive metadata labels for filtering, analysis, and domain-specific training. Dataset Information Total Chunks: 1,329 Format: jsonl-text Sorted: By document ID and chunk index (maintains document continuity) Source: Legal documents processed through enhanced parser with… See the full description on the dataset page: https://huggingface.co/datasets/rzeraat/legal-chunks-cpt.tabulartext-generation1K<n<10K0 likes17 downloads11mo agoHugging Face25NaruseShiroha /Genshin-CPT-Public License & Citation This dataset is a derivative work based on mrzjy/multimodal-genshin-impact. Original License: CC-BY-SA (Creative Commons Attribution-ShareAlike) Citation / Attribution: This processed dataset is derived from the work of mrzjy. As per the CC-BY-SA license, this derivative work is also released under the same license. Original Dataset Citation: mrzjy. (2023). multimodal-genshin-impact [Dataset]. Hugging Face.… See the full description on the dataset page: https://huggingface.co/datasets/NaruseShiroha/Genshin-CPT-Public.text10K<n<100K0 likes16 downloads9mo agoHugging Face26ananddey /asm-cpt-mixedtext1M<n<10M0 likes16 downloads3mo agoHugging Face27gitglubber /TL-001-CPT-Final-RTTtext10K<n<100K0 likes15 downloads4mo agoHugging Face28AdityaNarayan /HyperSwitch-Repo-CPT-Dataset-v2 Hyperswitch Rust Codebase Dataset A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models. 📊 Dataset Overview This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset-v2.tabulartext-generation10K<n<100K0 likes14 downloads10mo agoHugging Face29fenyo /H100-MonEspaceSante-CPT-corpus 🔧 Code & reproduction complète (scripts, RUNBOOK, reproduce.sh, conversations) : https://github.com/AlexandreFenyo/MonEspaceSante-H100-reproduction Mon espace santé — Corpus synthétique pour CPT (EntiGraph, FR) Corpus synthétique en français pour le continued pre-training (CPT) d'un LLM, destiné à injecter dans les poids les connaissances de la FAQ du service public « Mon espace santé ». Provenance (1-hop, anti model-collapse) Généré uniquement à partir des 88… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/H100-MonEspaceSante-CPT-corpus.texttext-generation10K<n<100K0 likes14 downloads4mo agoHugging Face30AdityaNarayan /HyperSwitch-Repo-CPT-Dataset Hyperswitch Rust Codebase Dataset A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models. 📊 Dataset Overview This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset.tabulartext-generation10K<n<100K1 likes13 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.