CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes232 downloads9mo agoHugging Face02FreedomIntelligence /HuatuoGPT2-Pretraining-Instruction HuatuoGPT2-Pretraining-Instruction-5200K Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT. This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible. Data Volume The following table details the volume and distribution of pre-training data for HuatuoGPT2: Data Source Data Volume Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.textquestion-answering1M<n<10M13 likes178 downloads2y agoHugging Face03windprak /steuerllm_pretraining_dataset SteuerLLM Pretraining Dataset Project page | Paper | GitHub Pretraining Dataset for German Tax Law filtered from FineWeb. This dataset was used for the continual pretraining stage of SteuerLLM, a specialized large language model for German tax law analysis. Dataset Description The SteuerLLM pretraining dataset is a domain-specific subset filtered from large-scale web corpora. It focuses on identifying and extracting tax-related content from German web data to adapt… See the full description on the dataset page: https://huggingface.co/datasets/windprak/steuerllm_pretraining_dataset.texttext-generation1M<n<10M1 likes113 downloads8mo agoHugging Face04lightman7 /tibetan-pretraining-corpus Tibetan Pre-training Corpus This dataset provides a comprehensive Tibetan language corpus developed for pre-training large language models. It contains carefully curated text from diverse sources to ensure broad coverage of the Tibetan language. Dataset Description The corpus consists of approximately 500,000 words of Tibetan text collected and processed specifically for large language model adaptation. The data has been sourced from: Tibetan Wikipedia content (70% of… See the full description on the dataset page: https://huggingface.co/datasets/lightman7/tibetan-pretraining-corpus.texttext-generation1M<n<10M0 likes19 downloads1y agoHugging Face05qingdu-giter /HuatuoGPT2-Pretraining-Instruction HuatuoGPT2-Pretraining-Instruction-5200K Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT. This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible. Data Volume The following table details the volume and distribution of pre-training data for HuatuoGPT2: Data Source Data Volume Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/qingdu-giter/HuatuoGPT2-Pretraining-Instruction.textquestion-answering1M<n<10M0 likes13 downloads6mo agoHugging Face06travisp83 /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/travisp83/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M0 likes8 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.