datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
continued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.EOS-Continued-Pretraining-Dataset
EOS Continued Pre-Training Dataset (Indonesia)
Deskripsi Dataset
EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia berskala besar (~214.2 Juta Token) yang dikurasi secara khusus untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM).
Dataset ini disusun sebagai bagian dari program Artificial Intelligence Talent Factory (AITF), kolaborasi antara Kementerian Komunikasi dan Digital (Komdigi) Republik Indonesia dan Universitas… See the full description on the dataset page: https://huggingface.co/datasets/aitf-komdigi/EOS-Continued-Pretraining-Dataset.EOS-Continued-Pretraining-Dataset
EOS Continued Pre-Training Dataset (Indonesia)
Deskripsi Dataset
EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia yang dikurasi untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM).
Tujuan utama dari dataset ini adalah untuk melakukan Domain Adaptation, yaitu meningkatkan kemampuan model dalam memahami konteks, terminologi, dan nuansa pada dua domain strategis di Indonesia:
Pengawasan Ruang Digital (PRD)
Digital Talent Pool… See the full description on the dataset page: https://huggingface.co/datasets/taqiyudinadn/EOS-Continued-Pretraining-Dataset.Continued-Pre-Training-Vocab-Paite
Paite Vocabulary — CPT Paragraph Text (vocab_paite_2025-12-13_paragraph.jsonl)
This file is continued pretraining (CPT) data: long plain-text sequences for causal language modeling. There is no instruction header—only a text field per line—so you can adapt token statistics and bilingual bridging patterns before instruction tuning.
Dataset composition
Construction: Built from the same Paite vocabulary sentence pairs as the SFT release. Each underlying example uses one of… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-Vocab-Paite.Continued-Pre-Training-CPT-Paite
Continued-Pre-Training-CPT-Paite (Master Collection)
This repository contains the unified, high-density raw text data used for the Continued Pre-Training (CPT) of the Sensix Paite models (Gemma-4-31B Master and Gemma-4-2B/5B Nitro).
The dataset is specifically designed to expand a base model's vocabulary and internalize Paite linguistic patterns, syntax, and tonal logic before moving to instruction fine-tuning (SFT).
Dataset Composition
This is a unified dataset… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-CPT-Paite.
