CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jiviteshjn /fineweb-edu-zh-chengyu-cpt Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.tabulartext-generation1M<n<10M1 likes1.8k downloads2mo agoHugging Face02proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes344 downloads5mo agoHugging Face03cptekur /pinchbench-clawd PinchBench Clawd Training Data Synthetic fine-tuning dataset for training an LLM to act as Clawd, an autonomous AI agent on the OpenClaw framework. Targets the PinchBench benchmark (23 tasks). Dataset Description Each example is a multi-turn conversation where Clawd uses tools (file I/O, web search, email, calendar, image generation, memory, etc.) to complete a real-world task. Generated using Claude via the Anthropic Batch API, scored by an LLM judge (1-5), and filtered… See the full description on the dataset page: https://huggingface.co/datasets/cptekur/pinchbench-clawd.texttext-generation1K<n<10K0 likes124 downloads6mo agoHugging Face04AmareshHebbar /cpt-coder-sft CPT / HCPCS Procedure Coder Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Procedure descriptions → correct CPT/HCPCS code with RVU data Why download this Build procedure coding assistants, verify CPT code assignments, or automate outpatient charge capture. Covers all specialties in the CMS PFS. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/cpt-coder-sft.texttext-generation10K<n<100K0 likes100 downloads3mo agoHugging Face05Chamaka8 /serendip-cpt-sinhala Serendib LLM CPT Sinhala Corpus A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for Continual Pre-Training (CPT) of large language models. This dataset was used to adapt Meta-LLaMA-3-8B to the Sinhala language domain as part of the Serendib LLM Honours Degree Research Project at the University of Central Lancashire (UCLan), 2025–2026. This is one of the largest openly published Sinhala NLP corpora available, containing 23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.texttext-generation10M<n<100M0 likes79 downloads6mo agoHugging Face06jiviteshjn /mc4-zh-idiom-cpt mC4 zh — Idiom-Tagged Continued-Pretraining Corpus A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in figurative language. Each document is natural web text (from the C4/mC4 zh subset) containing at least one culturally meaningful chengyu, with an appended knowledge block that lists every matched idiom together with its figurative meaning(s) and classical source citation. Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.tabulartext-generation1M<n<10M0 likes59 downloads2mo agoHugging Face07archit11 /cpt-dataset Hyperswitch CPT Dataset A comprehensive Continual Pre-Training (CPT) dataset for the Hyperswitch payment processing platform, combining documentation with actual code to build a "world model" understanding of the codebase. Dataset Description This dataset was created by mining the Hyperswitch repository and combining it with DeepWiki documentation. It teaches models: Repository Structure - Where different types of code live Concept-to-Code Mapping - How abstract concepts… See the full description on the dataset page: https://huggingface.co/datasets/archit11/cpt-dataset.texttext-generationn<1K0 likes58 downloads11mo agoHugging Face08fenyo /L40S-MonEspaceSante-CPT-corpus Mon Espace Santé — Corpus CPT synthétique Corpus synthétique de continued pre-training (CPT) généré à partir de 88 faits réels (paires Q/R « grounded ») de la FAQ du service public français Mon espace santé (fenyo/MonEspaceSante-FAQ-QA, source == "real"). But : injecter ces connaissances dans les poids d'un LLM sans RAG. Ce corpus sert au CPT décrit dans le modèle fenyo/L40S-Qwen3-8B-MonEspaceSante-CPT-SFT. Génération (résumé) Générateur : Qwen3-32B-FP8 servi par… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/L40S-MonEspaceSante-CPT-corpus.texttext-generation10K<n<100K0 likes28 downloads4mo agoHugging Face09aikathata /indonesian-dfk-cpt-dataset Indonesian DFK Domain CPT Corpus Dataset Description Dataset ini merupakan korpus teks Bahasa Indonesia untuk kebutuhan Continued Pre-Training atau CPT pada domain DFK, yaitu domain yang berkaitan dengan topik-topik yang sering menjadi sasaran disinformasi, fitnah, dan kebencian di Indonesia. Istilah DFK dalam dataset ini tidak berarti bahwa teks berisi disinformasi, fitnah, atau ujaran kebencian. DFK di sini merujuk pada domain atau topik yang sering menjadi sasaran DFK… See the full description on the dataset page: https://huggingface.co/datasets/aikathata/indonesian-dfk-cpt-dataset.texttext-generation1M<n<10M1 likes26 downloads4mo agoHugging Face10PoSTMEDIA /rosetta-ko-law-synth-cptgated rosetta-ko-law-synth-cpt Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Law Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-law-synth-cpt (this repo) continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-cpt.texttext-generation100K<n<1M0 likes22 downloads14d agoHugging Face11PoSTMEDIA /rosetta-ko-tourism-synth-cptgated rosetta-ko-tourism-synth-cpt Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Tourism Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-tourism-synth-cpt (this repo) continued-pretraining corpus (plain… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-cpt.texttext-generation10K<n<100K0 likes22 downloads14d agoHugging Face12rzeraat /legal-chunks-cpt Legal Document Chunks for Continued Pretraining This dataset contains 1329 legal document chunks extracted from various legal documents across multiple jurisdictions. Each chunk is enriched with comprehensive metadata labels for filtering, analysis, and domain-specific training. Dataset Information Total Chunks: 1,329 Format: jsonl-text Sorted: By document ID and chunk index (maintains document continuity) Source: Legal documents processed through enhanced parser with… See the full description on the dataset page: https://huggingface.co/datasets/rzeraat/legal-chunks-cpt.tabulartext-generation1K<n<10K0 likes18 downloads11mo agoHugging Face13fenyo /H100-MonEspaceSante-CPT-corpus 🔧 Code & reproduction complète (scripts, RUNBOOK, reproduce.sh, conversations) : https://github.com/AlexandreFenyo/MonEspaceSante-H100-reproduction Mon espace santé — Corpus synthétique pour CPT (EntiGraph, FR) Corpus synthétique en français pour le continued pre-training (CPT) d'un LLM, destiné à injecter dans les poids les connaissances de la FAQ du service public « Mon espace santé ». Provenance (1-hop, anti model-collapse) Généré uniquement à partir des 88… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/H100-MonEspaceSante-CPT-corpus.texttext-generation10K<n<100K0 likes14 downloads4mo agoHugging Face14AdityaNarayan /HyperSwitch-Repo-CPT-Dataset Hyperswitch Rust Codebase Dataset A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models. 📊 Dataset Overview This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset.tabulartext-generation10K<n<100K1 likes13 downloads11mo agoHugging Face15rzeraat /legal-chunks-cpt-test Legal Document Chunks for Continued Pretraining This dataset contains 1281 legal document chunks extracted from various legal documents across multiple jurisdictions. Each chunk is enriched with comprehensive metadata labels for filtering, analysis, and domain-specific training. Dataset Information Total Chunks: 1,281 Format: jsonl-text Sorted: By document ID and chunk index (maintains document continuity) Source: Legal documents processed through enhanced parser with… See the full description on the dataset page: https://huggingface.co/datasets/rzeraat/legal-chunks-cpt-test.texttext-generation1K<n<10K0 likes13 downloads11mo agoHugging Face16AdityaNarayan /HyperSwitch-Repo-CPT-Dataset-v2 Hyperswitch Rust Codebase Dataset A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models. 📊 Dataset Overview This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset-v2.tabulartext-generation10K<n<100K0 likes12 downloads11mo agoHugging Face17avemio /German-RAG-CPT-HESSIAN-AIgated German-RAG-CPT (Continued Pre-Training) Tasks Dataset German-RAG - German Retrieval Augmented Generation Dataset Summary The CPT Tasks Dataset is a comprehensive collection designed for continued pre-training of language models, focusing on three core competencies: context-based question answering, structured reasoning, and summarization. The dataset comprises approximately 620,000 examples, with 420,000 in German and 200,000 in English. Developed by Avemio AG… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-CPT-HESSIAN-AI.textquestion-answering100K<n<1M0 likes10 downloads2y agoHugging Face18sensix-zo /Continued-Pre-Training-CPT-Paitegated Continued-Pre-Training-CPT-Paite (Master Collection) This repository contains the unified, high-density raw text data used for the Continued Pre-Training (CPT) of the Sensix Paite models (Gemma-4-31B Master and Gemma-4-2B/5B Nitro). The dataset is specifically designed to expand a base model's vocabulary and internalize Paite linguistic patterns, syntax, and tonal logic before moving to instruction fine-tuning (SFT). Dataset Composition This is a unified dataset… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-CPT-Paite.texttext-generation1K<n<10K0 likes5 downloads5mo agoHugging Face19hicham-taoufik /gold-silver-mineral-process-cpt-candidates Gold/silver mineral-process CPT candidates English raw documents (text) about gold/silver and transferable hard-rock mineral processing. Source: BAAI/IndustryCorpus2_mining revision bf358a2f8105e4ac468141796e5a1a530685ae2e. English only. These are documents, not chat pairs. Configs Config Rows Notes default 59,749 all English bands english_high 18,864 publisher quality 4.00–4.59 english_middle 34,601 publisher quality 3.00–4.00 english_low 6,284… See the full description on the dataset page: https://huggingface.co/datasets/hicham-taoufik/gold-silver-mineral-process-cpt-candidates.tabulartext-generation100K<n<1M0 likes14h agoHugging Face20otisberg /jwiki_cpt Dataset Card for J Wiki Continuous Pretraining Markdown documents converted from the Jsoftware wiki for continuous pretraining (next-token prediction) of models that should read and write the J programming language. Each row is one wiki page. J session logs and scripts are fenced as j code blocks. Leading spaces, _ negatives (for example 3j_5, _1), and boxed-array characters are kept. Dataset Details Dataset Sources Wiki:… See the full description on the dataset page: https://huggingface.co/datasets/otisberg/jwiki_cpt.texttext-generation1K<n<10K0 likes20h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.