datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ettin-pretraining-data
Ettin Pre-training Data
Phase 1 of 3: Diverse pre-training data mixture (1.7T tokens) used to train the Ettin model suite.
This dataset contains the pre-training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.
📊 Data Composition
Data Source
Tokens (B)
Percentage
Description
DCLM
837.2
49.1%
High-quality web crawl data
CC Head
356.6… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-pretraining-data.Nemotron-Pretraining-Specialized-v1
Nemotron-Pre-Training-Dataset-v2.1
Dataset Description
The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.Nemotron-Pretraining-Code-v1
Nemotron-Pre-Training-Dataset-v1 Release
Data Overview
This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models.
This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1.Nemotron-Pretraining-Code-v2
Nemotron-Pre-Training-Dataset-v2.1
Dataset Description
The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2.carbon-pretraining-corpus
🧬 Carbon Pretraining Corpus
Description
173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model.
This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species.
Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.Nemotron-Pretraining-Specialized-v1.1
Nemotron-Pretraining-Specialized-v1.1
Dataset Description:
The Nemotron-Pretraining-Specialized-v1.1 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities in code concepts and algorithms, formal logic, economics, and multiple choice questions. The code concepts dataset is an instance of a general… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.1.Nemotron-Pretraining-Specialized-v1.2
Nemotron-Pretraining-Specialized-v1.2
Dataset Description:
The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions.
Note: These are new datasets, not replacements.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.2.Nemotron-Pretraining-SFT-v1
Nemotron-Pre-Training-Dataset-v1 Release
Data Overview
This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models.
This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1.propagator-multimodal-pretraining-data
Propagator Multimodal Pretraining Data
This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format.
This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout.
Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.Nemotron-Pretraining-Legal-v1
Nemotron-Pretraining-Legal-v1
Dataset Description:
The Nemotron-Pretraining-Legal-v1 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets intended to improve the legal capabilities of LLMs. In one ablation, adding these datasets to Nemotron 3 Nano pretraining boosted a proxy LegalBench average accuracy from 64.6 to 74.7.
This… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Legal-v1.Nemotron-Pretraining-Code-v3
Nemotron-Pretraining-Code-v3
Dataset Description:
The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs.
The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v3.TR-HASH-Pretraining-125B-Agentic-32K
TR-HASH Pretraining 125B — Agentic 32K
Private, source-curated pretraining artifact for the TR-HASH Agentic 32K model
line. It contains 125B packed token exposures: 75B foundation and 50B
agentic/procedural content.
The corpus uses the immutable, validated 32,000-ID revision of
AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic. It is not compatible with the
older TR-HASH 32K tokenizer.
Composition
Bucket
Tokens
Purpose
Foundation
75B
English and French… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/TR-HASH-Pretraining-125B-Agentic-32K.KcBERT_Pre-Training_Corpus
KcBERT Pre-Training Corpus (Korean News Comments)
KcBERT
beomi/kcbert-base
Github KcBERT Repo: https://github.com/Beomi/KcBERTKcBERT is Korean Comments BERT pretrained on this Corpus set.(You can use it via Huggingface's Transformers library!)
This Kaggle Dataset contains CLEANED dataset preprocessed with the code below.
import re
import emoji
from soynlp.normalizer import repeat_normalize
emojis = ''.join(emoji.UNICODE_EMOJI.keys())
pattern = re.compile(f'[^ .… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/KcBERT_Pre-Training_Corpus.Nemotron-Pretraining-Specialized-v1
Nemotron-Pre-Training-Dataset-v2.1
Dataset Description
The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Nemotron-Pretraining-Specialized-v1.stratified-kmeans-diverse-pretraining-100K-1M
Stratified K-Means Diverse Pre-Training Dataset (100K-1M)
A carefully balanced subset combining FineWeb-Edu and Proof-Pile-2, featuring embedding-based k-means sampling to ensure diverse representation across educational and mathematical/scientific content at multiple scales.
👥 Follow the Authors
Aman Priyanshu
Supriti Vijay
Overview
This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales, combining high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-pretraining-100K-1M.Nemotron-Pretraining-Code-v3
Nemotron-Pretraining-Code-v3
Dataset Description:
The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs.
The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/TerraBytes/Nemotron-Pretraining-Code-v3.pretraining-corpus
🧠 Kiy-K Synthetic Pretraining Corpus
Author: Khoi K. (@Kiy-K)License: Apache 2.0Last Updated: 2025-10-30
📘 Overview
The Kiy-K Synthetic Pretraining Corpus is a large-scale collection of synthetically generated English text designed for language model pretraining and instruction-tuning research.
All data is synthetic, created using open-source large language models such as GPT-OSS, NVIDIA Nemotron, and DeepSeek, under full control of the author.No real user… See the full description on the dataset page: https://huggingface.co/datasets/Kiy-K/pretraining-corpus.Nemotron-Pretraining-Specialized-v1.2
Nemotron-Pretraining-Specialized-v1.2
Dataset Description:
The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions.
Note: These are new datasets, not replacements.… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Nemotron-Pretraining-Specialized-v1.2.Nemotron-Pretraining-Code-v3
Nemotron-Pretraining-Code-v3
Dataset Description:
The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs.
The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/WillowVoiceAI/Nemotron-Pretraining-Code-v3.amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic.
You can load the dataset as follows
from datasets import load_dataset
ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus")
curated-pretraining
Curated Pre-Tokenized Training Dataset
Pre-tokenized training data for a 500M parameter LLaMA-style model optimized
for structured output tasks (JSON generation, function calling, schema compliance).
Format
Binary shards of packed uint16 token sequences. Each shard has a 16-byte header
followed by contiguous sequences of 2048 tokens.
Tokenizer: EleutherAI/gpt-neox-20b (vocab size: 50,304)
Context length: 2048
Total tokens: 19.86B
Total sequences: 9,696,327
Shards: 74… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/curated-pretraining.tiny-slm-pretraining-corpus
🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB)
A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures.
100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders.
📊 Dataset Statistics
Total Documents: 20,066,075
Train: 19,663,898
Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.1gpu-llm-pretraining-corpus-15b-en-it-code
1GPU LLM Pretraining Corpus 15B EN-IT-CODE
1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family.
It was built for training language models from scratch on a mixture of English, Italian and source code.
This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.nemotron-pretraining-specialized-collection
Nemotron Pretraining Specialized Collection
This repository is a convenience collection of the NVIDIA Nemotron Pretraining Specialized releases. It preserves each original subset as a separate Hugging Face configuration, so consumers can select a single domain-focused subset without visiting multiple source repositories.
Contents
Source release
Configurations included
nvidia/Nemotron-Pretraining-Specialized-v1
Wiki Rewrite, Math Textbooks, STEM SFT… See the full description on the dataset page: https://huggingface.co/datasets/CrowdMind/nemotron-pretraining-specialized-collection.swahili-pretraining-corpus
Swahili Pretraining Corpus (tokenized)
The pre-training corpus used to train Benjamin-png/swahili-gpt-71m
from scratch — ~2.15 billion tokens of real Swahili text, already tokenized
and ready for language-model training.
Contents
File
What it is
train.bin
Training tokens — a flat uint16 array of token IDs (~2.15B tokens, ~4.3 GB)
val.bin
Held-out validation tokens (same format)
swahili_tokenizer.model
The SentencePiece tokenizer (32k vocab… See the full description on the dataset page: https://huggingface.co/datasets/Benjamin-png/swahili-pretraining-corpus.HuatuoGPT2-Pretraining-Instruction
HuatuoGPT2-Pretraining-Instruction-5200K
Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT.
This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible.
Data Volume
The following table details the volume and distribution of pre-training data for HuatuoGPT2:
Data Source
Data Volume
Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.Nemotron-Pretraining-Code-v3
Nemotron-Pretraining-Code-v3
Dataset Description:
The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs.
The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/P3rc3us/Nemotron-Pretraining-Code-v3.Nemotron-Pretraining-Specialized-v1
Nemotron-Pre-Training-Dataset-v2.1
Dataset Description
The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/semran1/Nemotron-Pretraining-Specialized-v1.
