continued-pretraining
100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-Q4_K_M-GGUFtest_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bittest_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-Q8_0-GGUF
continued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.EOS-Continued-Pretraining-Dataset
EOS Continued Pre-Training Dataset (Indonesia)
Deskripsi Dataset
EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia berskala besar (~214.2 Juta Token) yang dikurasi secara khusus untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM).
Dataset ini disusun sebagai bagian dari program Artificial Intelligence Talent Factory (AITF), kolaborasi antara Kementerian Komunikasi dan Digital (Komdigi) Republik Indonesia dan Universitas… See the full description on the dataset page: https://huggingface.co/datasets/aitf-komdigi/EOS-Continued-Pretraining-Dataset.EOS-Continued-Pretraining-Dataset
EOS Continued Pre-Training Dataset (Indonesia)
Deskripsi Dataset
EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia yang dikurasi untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM).
Tujuan utama dari dataset ini adalah untuk melakukan Domain Adaptation, yaitu meningkatkan kemampuan model dalam memahami konteks, terminologi, dan nuansa pada dua domain strategis di Indonesia:
Pengawasan Ruang Digital (PRD)
Digital Talent Pool… See the full description on the dataset page: https://huggingface.co/datasets/taqiyudinadn/EOS-Continued-Pretraining-Dataset.creole-text-continued-pretraining
