CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-paws /continued-pretraining-llama-format Open Paws Continued Pretraining Llama Format Overview This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Specialized Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.texttext-generation10K<n<100K2 likes61 downloads1y agoHugging Face02open-llm-leaderboard /FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face03open-llm-leaderboard /FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face04LexiconShiftInnovations /Continued_Pretrained_Dataset_Dental_Sinhalatext1K<n<10K0 likes30 downloads2y agoHugging Face05aitf-komdigi /EOS-Continued-Pretraining-Dataset EOS Continued Pre-Training Dataset (Indonesia) Deskripsi Dataset EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia berskala besar (~214.2 Juta Token) yang dikurasi secara khusus untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM). Dataset ini disusun sebagai bagian dari program Artificial Intelligence Talent Factory (AITF), kolaborasi antara Kementerian Komunikasi dan Digital (Komdigi) Republik Indonesia dan Universitas… See the full description on the dataset page: https://huggingface.co/datasets/aitf-komdigi/EOS-Continued-Pretraining-Dataset.texttext-generation100K<n<1M1 likes24 downloads9mo agoHugging Face06jsbeaudry /creole-text-continued-pretrainingtext10K<n<100K0 likes14 downloads1y agoHugging Face07taqiyudinadn /EOS-Continued-Pretraining-Dataset EOS Continued Pre-Training Dataset (Indonesia) Deskripsi Dataset EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia yang dikurasi untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM). Tujuan utama dari dataset ini adalah untuk melakukan Domain Adaptation, yaitu meningkatkan kemampuan model dalam memahami konteks, terminologi, dan nuansa pada dua domain strategis di Indonesia: Pengawasan Ruang Digital (PRD) Digital Talent Pool… See the full description on the dataset page: https://huggingface.co/datasets/taqiyudinadn/EOS-Continued-Pretraining-Dataset.texttext-generation100K<n<1M0 likes11 downloads9mo agoHugging Face08open-llm-leaderboard /FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face09open-llm-leaderboard /FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face10sensix-zo /Continued-Pre-Training-Vocab-Paitegated Paite Vocabulary — CPT Paragraph Text (vocab_paite_2025-12-13_paragraph.jsonl) This file is continued pretraining (CPT) data: long plain-text sequences for causal language modeling. There is no instruction header—only a text field per line—so you can adapt token statistics and bilingual bridging patterns before instruction tuning. Dataset composition Construction: Built from the same Paite vocabulary sentence pairs as the SFT release. Each underlying example uses one of… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-Vocab-Paite.texttext-generation10K<n<100K0 likes7 downloads5mo agoHugging Face11open-llm-leaderboard /FlofloB__test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__test_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face12open-llm-leaderboard /FlofloB__83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__83k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face13Maverfrick /Continued-Pretraining-Vietnamese-TCMtext1K<n<10K0 likes5 downloads7mo agoHugging Face14sensix-zo /Continued-Pre-Training-CPT-Paitegated Continued-Pre-Training-CPT-Paite (Master Collection) This repository contains the unified, high-density raw text data used for the Continued Pre-Training (CPT) of the Sensix Paite models (Gemma-4-31B Master and Gemma-4-2B/5B Nitro). The dataset is specifically designed to expand a base model's vocabulary and internalize Paite linguistic patterns, syntax, and tonal logic before moving to instruction fine-tuning (SFT). Dataset Composition This is a unified dataset… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-CPT-Paite.texttext-generation1K<n<10K0 likes5 downloads5mo agoHugging Face15Inabia-AI /continued-pretrainingtext10K<n<100K0 likes2 downloads2y agoHugging Face16Inabia-AI /continued-pretraining-jsonltext10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.