CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jiviteshjn /fineweb-edu-zh-chengyu-cpt Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.tabulartext-generation1M<n<10M1 likes1.9k downloads2mo agoHugging Face02proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes338 downloads5mo agoHugging Face03atenareply /asterion-cpt-corpus Asterion CPT Source Corpus What's inside — 297,185 synthetic technical documents about the fictional Asterion Space Operations fleet (24 satellites: EO/COMM/SCI/TD): 7 doc_types × 12 topics, ~2.5 GB of text, ~1.62B Gemma-4 tokens (measured mean 5,413 tokens/doc on a stratified 2k sample, 2026-07-03). Where it comes from — Fully synthetic — generated incrementally in 1,000-doc shards by the Asterion corpus generation pipeline (see noval-corp/docs/asterion-corpus-plan.md) from a… See the full description on the dataset page: https://huggingface.co/datasets/atenareply/asterion-cpt-corpus.tabulartext-generation100K<n<1M0 likes314 downloads3mo agoHugging Face04DimitarV /bulgarian-medical-cpt-100m Bulgarian text for MOSS continued pretraining Exactly 100 million training tokens: 10M medical and 90M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 10,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.tabulartext-generation100K<n<1M0 likes85 downloads13d agoHugging Face05louisbrulenaudet /Romulus-cpt-fr Romulus, continually pre-trained models for French law. Romulus is a series of continually pre-trained models enriched in French law and intended to serve as the basis for a fine-tuning process on labeled data. Please note that these models have not been aligned for the production of usable text as they stand, and will certainly need to be fine-tuned for the desired tasks in order to produce satisfactory results. The training corpus is made up of around 34,864,949 tokens… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/Romulus-cpt-fr.tabulartext-generation100K<n<1M5 likes69 downloads2y agoHugging Face06jiviteshjn /mc4-zh-idiom-cpt mC4 zh — Idiom-Tagged Continued-Pretraining Corpus A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in figurative language. Each document is natural web text (from the C4/mC4 zh subset) containing at least one culturally meaningful chengyu, with an appended knowledge block that lists every matched idiom together with its figurative meaning(s) and classical source citation. Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.tabulartext-generation1M<n<10M0 likes69 downloads2mo agoHugging Face07DimitarV /bulgarian-medical-cpt-10m Bulgarian text for MOSS continued pretraining Exactly 10 million training tokens: 3M medical and 7M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 3,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-10m.tabulartext-generation10K<n<100K0 likes64 downloads13d agoHugging Face08jhdlee /wiki-events-cpt Wikipedia Events CPT jhdlee/wiki-events-cpt is a public research dataset of 150 selected English Wikipedia articles with compact metadata for continual pretraining (CPT). Split Articles Event window (end exclusive) cohort_a 75 2023-01-01 to 2024-10-01 cohort_b 75 2024-10-01 to 2025-09-01 Each cohort has 25 articles per topic: natural_hazards, elections, and sports. Cohorts group events by their reviewed whole-occurrence intervals; they are not… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/wiki-events-cpt.tabulartext-generation10K<n<100K0 likes60 downloads15d agoHugging Face09oliveirabruno01 /ptbr-creative-cpt-qwen35-08b-v02 PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2 This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments. It is not the canonical text corpus. Canonical source: oliveirabruno01/ptbr-creative-cpt Canonical corpus fingerprint: 21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840 Identity Model/tokenizer: Qwen/Qwen3.5-0.8B-Base Context length: 2048 Data-prep version: v0.2 Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.tabulartext-generation1K<n<10K0 likes56 downloads2d agoHugging Face10oliveirabruno01 /ptbr-creative-cpt PT-BR Creative Corpus v0.1.0 A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research. Status This is the canonical corpus freeze, not a final model-specific training build. Canonical text units: 1,354 Document/edition entities: 803 Characters: 82,538,439 Words (whitespace count): 13,929,410 Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config. The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.tabulartext-generation1K<n<10K0 likes54 downloads2d agoHugging Face11rzeraat /legal-chunks-cpt Legal Document Chunks for Continued Pretraining This dataset contains 1329 legal document chunks extracted from various legal documents across multiple jurisdictions. Each chunk is enriched with comprehensive metadata labels for filtering, analysis, and domain-specific training. Dataset Information Total Chunks: 1,329 Format: jsonl-text Sorted: By document ID and chunk index (maintains document continuity) Source: Legal documents processed through enhanced parser with… See the full description on the dataset page: https://huggingface.co/datasets/rzeraat/legal-chunks-cpt.tabulartext-generation1K<n<10K0 likes17 downloads11mo agoHugging Face12AdityaNarayan /HyperSwitch-Repo-CPT-Dataset-v2 Hyperswitch Rust Codebase Dataset A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models. 📊 Dataset Overview This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset-v2.tabulartext-generation10K<n<100K0 likes14 downloads10mo agoHugging Face13AdityaNarayan /HyperSwitch-Repo-CPT-Dataset Hyperswitch Rust Codebase Dataset A comprehensive dataset extracted from the Hyperswitch open-source payment processing platform, containing 16,731 code samples across 37 modules with 6.99M tokens for training Rust code understanding and generation models. 📊 Dataset Overview This dataset provides both file-level and granular code samples from Hyperswitch, a modern payment switch written in Rust. It's designed for training code models to understand payment processing… See the full description on the dataset page: https://huggingface.co/datasets/AdityaNarayan/HyperSwitch-Repo-CPT-Dataset.tabulartext-generation10K<n<100K1 likes13 downloads11mo agoHugging Face14costadev00 /stf-acordaos-cpt-2048gated STF Acordaos CPT 2048 Dataset de acórdãos do STF preparado para continuous pre-training (CPT), com foco em textos jurídicos e preservação da cauda final dos documentos longos (parte 2, parte 3, etc.). Origem dos dados Fonte pública: https://dadosabertos.c3sl.ufpr.br/acordaos/json/ Arquivos de origem utilizados: DocumentosAcordaos.json AcordaosVotos.json AcordaosRelatorios.json Construção Fonte canônica: DocumentosAcordaos.json Labels incluídos: integra, voto… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/stf-acordaos-cpt-2048.tabulartext-generation100K<n<1M0 likes9 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.