CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jhu-clsp /ettin-pretraining-data Ettin Pre-training Data Phase 1 of 3: Diverse pre-training data mixture (1.7T tokens) used to train the Ettin model suite. This dataset contains the pre-training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source Tokens (B) Percentage Description DCLM 837.2 49.1% High-quality web crawl data CC Head 356.6… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-pretraining-data.text-generation10 likes152k downloads1y agoHugging Face02nvidia /Nemotron-Pretraining-Specialized-v1 Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.texttext-generation10M<n<100M87 likes5.9k downloads9mo agoHugging Face03nvidia /Nemotron-Pretraining-Code-v1gated Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1.texttext-generation100M<n<1B78 likes4.3k downloads9mo agoHugging Face04nvidia /Nemotron-Pretraining-Code-v2gated Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2.texttext-generation100M<n<1B135 likes4.2k downloads9mo agoHugging Face05HuggingFaceBio /carbon-pretraining-corpus 🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.tabulartext-generation100M<n<1B30 likes3.6k downloads3mo agoHugging Face06nvidia /Nemotron-Pretraining-Specialized-v1.1 Nemotron-Pretraining-Specialized-v1.1 Dataset Description: The Nemotron-Pretraining-Specialized-v1.1 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities in code concepts and algorithms, formal logic, economics, and multiple choice questions. The code concepts dataset is an instance of a general… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.1.texttext-generation10M<n<100M46 likes2.7k downloads7mo agoHugging Face07nvidia /Nemotron-Pretraining-Specialized-v1.2 Nemotron-Pretraining-Specialized-v1.2 Dataset Description: The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions. Note: These are new datasets, not replacements.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.2.texttext-generation100M<n<1B16 likes2.7k downloads4mo agoHugging Face08nvidia /Nemotron-Pretraining-SFT-v1gated Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1.texttext-generation100M<n<1B73 likes2.1k downloads9mo agoHugging Face09ken-sungmin /propagator-multimodal-pretraining-data Propagator Multimodal Pretraining Data This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format. This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout. Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.texttext-generation0 likes1.4k downloads3mo agoHugging Face10nvidia /Nemotron-Pretraining-Legal-v1 Nemotron-Pretraining-Legal-v1 Dataset Description: The Nemotron-Pretraining-Legal-v1 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets intended to improve the legal capabilities of LLMs. In one ablation, adding these datasets to Nemotron 3 Nano pretraining boosted a proxy LegalBench average accuracy from 64.6 to 74.7. This… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Legal-v1.texttext-generation1M<n<10M25 likes1k downloads4mo agoHugging Face11nvidia /Nemotron-Pretraining-Code-v3 Nemotron-Pretraining-Code-v3 Dataset Description: The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs. The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v3.texttext-generation100M<n<1B73 likes1k downloads4mo agoHugging Face12AETHORIA-AI /TR-HASH-Pretraining-125B-Agentic-32K TR-HASH Pretraining 125B — Agentic 32K Private, source-curated pretraining artifact for the TR-HASH Agentic 32K model line. It contains 125B packed token exposures: 75B foundation and 50B agentic/procedural content. The corpus uses the immutable, validated 32,000-ID revision of AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic. It is not compatible with the older TR-HASH 32K tokenizer. Composition Bucket Tokens Purpose Foundation 75B English and French… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/TR-HASH-Pretraining-125B-Agentic-32K.text-generation0 likes754 downloads24d agoHugging Face13Bingsu /KcBERT_Pre-Training_Corpus KcBERT Pre-Training Corpus (Korean News Comments) KcBERT beomi/kcbert-base Github KcBERT Repo: https://github.com/Beomi/KcBERTKcBERT is Korean Comments BERT pretrained on this Corpus set.(You can use it via Huggingface's Transformers library!) This Kaggle Dataset contains CLEANED dataset preprocessed with the code below. import re import emoji from soynlp.normalizer import repeat_normalize emojis = ''.join(emoji.UNICODE_EMOJI.keys()) pattern = re.compile(f'[^ .… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/KcBERT_Pre-Training_Corpus.textfill-mask10M<n<100M1 likes643 downloads4y agoHugging Face14erenyeager-1 /Nemotron-Pretraining-Specialized-v1 Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Nemotron-Pretraining-Specialized-v1.texttext-generation10M<n<100M0 likes611 downloads1mo agoHugging Face15AmanPriyanshu /stratified-kmeans-diverse-pretraining-100K-1M Stratified K-Means Diverse Pre-Training Dataset (100K-1M) A carefully balanced subset combining FineWeb-Edu and Proof-Pile-2, featuring embedding-based k-means sampling to ensure diverse representation across educational and mathematical/scientific content at multiple scales. 👥 Follow the Authors Aman Priyanshu Supriti Vijay Overview This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales, combining high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-pretraining-100K-1M.texttext-generation1M<n<10M1 likes515 downloads1y agoHugging Face16TerraBytes /Nemotron-Pretraining-Code-v3 Nemotron-Pretraining-Code-v3 Dataset Description: The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs. The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/TerraBytes/Nemotron-Pretraining-Code-v3.texttext-generation100M<n<1B0 likes467 downloads3mo agoHugging Face17Kiy-K /pretraining-corpus 🧠 Kiy-K Synthetic Pretraining Corpus Author: Khoi K. (@Kiy-K)License: Apache 2.0Last Updated: 2025-10-30 📘 Overview The Kiy-K Synthetic Pretraining Corpus is a large-scale collection of synthetically generated English text designed for language model pretraining and instruction-tuning research. All data is synthetic, created using open-source large language models such as GPT-OSS, NVIDIA Nemotron, and DeepSeek, under full control of the author.No real user… See the full description on the dataset page: https://huggingface.co/datasets/Kiy-K/pretraining-corpus.tabulartext-generation10K<n<100K3 likes386 downloads10mo agoHugging Face18erenyeager-1 /Nemotron-Pretraining-Specialized-v1.2 Nemotron-Pretraining-Specialized-v1.2 Dataset Description: The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions. Note: These are new datasets, not replacements.… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Nemotron-Pretraining-Specialized-v1.2.texttext-generation100M<n<1B0 likes342 downloads1mo agoHugging Face19WillowVoiceAI /Nemotron-Pretraining-Code-v3 Nemotron-Pretraining-Code-v3 Dataset Description: The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs. The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/WillowVoiceAI/Nemotron-Pretraining-Code-v3.texttext-generation100M<n<1B0 likes315 downloads3mo agoHugging Face20yordanoswuletaw /amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic. You can load the dataset as follows from datasets import load_dataset ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus") texttext-generation100M<n<1B4 likes296 downloads2y agoHugging Face21mdonigian /curated-pretraining Curated Pre-Tokenized Training Dataset Pre-tokenized training data for a 500M parameter LLaMA-style model optimized for structured output tasks (JSON generation, function calling, schema compliance). Format Binary shards of packed uint16 token sequences. Each shard has a 16-byte header followed by contiguous sequences of 2048 tokens. Tokenizer: EleutherAI/gpt-neox-20b (vocab size: 50,304) Context length: 2048 Total tokens: 19.86B Total sequences: 9,696,327 Shards: 74… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/curated-pretraining.text-generation0 likes292 downloads7mo agoHugging Face22JustACluelessKidAtSchool /tiny-slm-pretraining-corpus 🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB) A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures. 100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders. 📊 Dataset Statistics Total Documents: 20,066,075 Train: 19,663,898 Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tabulartext-generation10M<n<100M0 likes224 downloads1mo agoHugging Face23visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes209 downloads9mo agoHugging Face24nazdef /1gpu-llm-pretraining-corpus-15b-en-it-code 1GPU LLM Pretraining Corpus 15B EN-IT-CODE 1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family. It was built for training language models from scratch on a mixture of English, Italian and source code. This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.text-generation0 likes184 downloads4d agoHugging Face25CrowdMind /nemotron-pretraining-specialized-collection Nemotron Pretraining Specialized Collection This repository is a convenience collection of the NVIDIA Nemotron Pretraining Specialized releases. It preserves each original subset as a separate Hugging Face configuration, so consumers can select a single domain-focused subset without visiting multiple source repositories. Contents Source release Configurations included nvidia/Nemotron-Pretraining-Specialized-v1 Wiki Rewrite, Math Textbooks, STEM SFT… See the full description on the dataset page: https://huggingface.co/datasets/CrowdMind/nemotron-pretraining-specialized-collection.texttext-generation100M<n<1B0 likes182 downloads1d agoHugging Face26Benjamin-png /swahili-pretraining-corpus Swahili Pretraining Corpus (tokenized) The pre-training corpus used to train Benjamin-png/swahili-gpt-71m from scratch — ~2.15 billion tokens of real Swahili text, already tokenized and ready for language-model training. Contents File What it is train.bin Training tokens — a flat uint16 array of token IDs (~2.15B tokens, ~4.3 GB) val.bin Held-out validation tokens (same format) swahili_tokenizer.model The SentencePiece tokenizer (32k vocab… See the full description on the dataset page: https://huggingface.co/datasets/Benjamin-png/swahili-pretraining-corpus.text-generation1B<n<10B0 likes173 downloads3mo agoHugging Face27FreedomIntelligence /HuatuoGPT2-Pretraining-Instruction HuatuoGPT2-Pretraining-Instruction-5200K Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT. This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible. Data Volume The following table details the volume and distribution of pre-training data for HuatuoGPT2: Data Source Data Volume Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.textquestion-answering1M<n<10M13 likes164 downloads2y agoHugging Face28lapa-llm /pretraining-high-quality Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness of the text lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.tabulartext-generation10M<n<100M0 likes139 downloads10mo agoHugging Face29P3rc3us /Nemotron-Pretraining-Code-v3 Nemotron-Pretraining-Code-v3 Dataset Description: The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs. The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/P3rc3us/Nemotron-Pretraining-Code-v3.texttext-generation100M<n<1B0 likes135 downloads26d agoHugging Face30semran1 /Nemotron-Pretraining-Specialized-v1 Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/semran1/Nemotron-Pretraining-Specialized-v1.texttext-generation10M<n<100M2 likes124 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.