CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes338 downloads5mo agoHugging Face02jhdlee /wiki-events-cpt Wikipedia Events CPT jhdlee/wiki-events-cpt is a public research dataset of 150 selected English Wikipedia articles with compact metadata for continual pretraining (CPT). Split Articles Event window (end exclusive) cohort_a 75 2023-01-01 to 2024-10-01 cohort_b 75 2024-10-01 to 2025-09-01 Each cohort has 25 articles per topic: natural_hazards, elections, and sports. Cohorts group events by their reviewed whole-occurrence intervals; they are not… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/wiki-events-cpt.tabulartext-generation10K<n<100K0 likes60 downloads15d agoHugging Face03archit11 /cpt-dataset Hyperswitch CPT Dataset A comprehensive Continual Pre-Training (CPT) dataset for the Hyperswitch payment processing platform, combining documentation with actual code to build a "world model" understanding of the codebase. Dataset Description This dataset was created by mining the Hyperswitch repository and combining it with DeepWiki documentation. It teaches models: Repository Structure - Where different types of code live Concept-to-Code Mapping - How abstract concepts… See the full description on the dataset page: https://huggingface.co/datasets/archit11/cpt-dataset.texttext-generationn<1K0 likes59 downloads11mo agoHugging Face04william-0g /MSA-cpt-100b MSA CPT corpus (cpt_100b) Continual-pre-training corpus for Memory Sparse Attention (MSA) / Generative Retrieval on a Qwen3.5 backbone. Built 2026-06-29 by aggregating and within-source-deduplicating 45 public retrieval / QA datasets into a unified query↔document contract. Layout corpus/ # document side (.jsonl.gz shards) train/ # query↔positive-doc training pairs (.jsonl.gz shards) manifest.json Stats (from manifest.json) field value… See the full description on the dataset page: https://huggingface.co/datasets/william-0g/MSA-cpt-100b.text-retrieval0 likes30 downloads2mo agoHugging Face05avemio /German-RAG-CPT-HESSIAN-AIgated German-RAG-CPT (Continued Pre-Training) Tasks Dataset German-RAG - German Retrieval Augmented Generation Dataset Summary The CPT Tasks Dataset is a comprehensive collection designed for continued pre-training of language models, focusing on three core competencies: context-based question answering, structured reasoning, and summarization. The dataset comprises approximately 620,000 examples, with 420,000 in German and 200,000 in English. Developed by Avemio AG… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-CPT-HESSIAN-AI.textquestion-answering100K<n<1M0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.