datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Continued-Pre-Training-Vocab-Paite
Paite Vocabulary — CPT Paragraph Text (vocab_paite_2025-12-13_paragraph.jsonl)
This file is continued pretraining (CPT) data: long plain-text sequences for causal language modeling. There is no instruction header—only a text field per line—so you can adapt token statistics and bilingual bridging patterns before instruction tuning.
Dataset composition
Construction: Built from the same Paite vocabulary sentence pairs as the SFT release. Each underlying example uses one of… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-Vocab-Paite.Continued-Pre-Training-CPT-Paite
Continued-Pre-Training-CPT-Paite (Master Collection)
This repository contains the unified, high-density raw text data used for the Continued Pre-Training (CPT) of the Sensix Paite models (Gemma-4-31B Master and Gemma-4-2B/5B Nitro).
The dataset is specifically designed to expand a base model's vocabulary and internalize Paite linguistic patterns, syntax, and tonal logic before moving to instruction fine-tuning (SFT).
Dataset Composition
This is a unified dataset… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-CPT-Paite.
