datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.HuatuoGPT2-Pretraining-Instruction
HuatuoGPT2-Pretraining-Instruction-5200K
Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT.
This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible.
Data Volume
The following table details the volume and distribution of pre-training data for HuatuoGPT2:
Data Source
Data Volume
Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.steuerllm_pretraining_dataset
SteuerLLM Pretraining Dataset
Project page | Paper | GitHub
Pretraining Dataset for German Tax Law filtered from FineWeb. This dataset was used for the continual pretraining stage of SteuerLLM, a specialized large language model for German tax law analysis.
Dataset Description
The SteuerLLM pretraining dataset is a domain-specific subset filtered from large-scale web corpora. It focuses on identifying and extracting tax-related content from German web data to adapt… See the full description on the dataset page: https://huggingface.co/datasets/windprak/steuerllm_pretraining_dataset.tibetan-pretraining-corpus
Tibetan Pre-training Corpus
This dataset provides a comprehensive Tibetan language corpus developed for pre-training large language models. It contains carefully curated text from diverse sources to ensure broad coverage of the Tibetan language.
Dataset Description
The corpus consists of approximately 500,000 words of Tibetan text collected and processed specifically for large language model adaptation. The data has been sourced from:
Tibetan Wikipedia content (70% of… See the full description on the dataset page: https://huggingface.co/datasets/lightman7/tibetan-pretraining-corpus.3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM
DATASET NAME: KS-LIT-3M
Kashmiri Pretraining Dataset
This repository hosts a meticulously processed Kashmiri text dataset, specifically designed for pretraining Large Language Models (LLMs) from scratch. The dataset has undergone extensive cleaning and preprocessing to ensure high quality and suitability for robust model training.
Dataset Description
This dataset consists of a continuous one stream of Kashmiri text, cleaned to remove English… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/3.1Million_KASHMIRI_text_Pre_training_Dataset_for_LLM_2026_by_HNM.HuatuoGPT2-Pretraining-Instruction
HuatuoGPT2-Pretraining-Instruction-5200K
Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT.
This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible.
Data Volume
The following table details the volume and distribution of pre-training data for HuatuoGPT2:
Data Source
Data Volume
Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/qingdu-giter/HuatuoGPT2-Pretraining-Instruction.agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/travisp83/agentic-llm-pretraining-1.7b.Continued-Pre-Training-Vocab-Paite
Paite Vocabulary — CPT Paragraph Text (vocab_paite_2025-12-13_paragraph.jsonl)
This file is continued pretraining (CPT) data: long plain-text sequences for causal language modeling. There is no instruction header—only a text field per line—so you can adapt token statistics and bilingual bridging patterns before instruction tuning.
Dataset composition
Construction: Built from the same Paite vocabulary sentence pairs as the SFT release. Each underlying example uses one of… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-Vocab-Paite.Continued-Pre-Training-CPT-Paite
Continued-Pre-Training-CPT-Paite (Master Collection)
This repository contains the unified, high-density raw text data used for the Continued Pre-Training (CPT) of the Sensix Paite models (Gemma-4-31B Master and Gemma-4-2B/5B Nitro).
The dataset is specifically designed to expand a base model's vocabulary and internalize Paite linguistic patterns, syntax, and tonal logic before moving to instruction fine-tuning (SFT).
Dataset Composition
This is a unified dataset… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-CPT-Paite.
