erenyeager-1/Nemotron-Pretraining-Specialized-v1.2
Nemotron-Pretraining-Specialized-v1.2 Dataset Description: The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions. Note: These are new datasets, not replacements.… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Nemotron-Pretraining-Specialized-v1.2.
Nemotron-Pretraining-Specialized-v1.2
Dataset Description:
The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions.
Note: These are new datasets, not replacements. They are meant to be used together with the previously released datasets Nemotron-Pretraining-Specialized-v1.1 and Nemotron-Pretraining-Specialized-v1.
This dataset is ready for commercial use.
Dataset Details:
For more details, please see the NVIDIA Nemotron 3 Ultra tech report.
This dataset has the following subsets:
- Nemotron-Pretraining-Fact-Seeking: Fact-seeking questions generated from Finewiki. In an ablation on an intermediate Nemotron 3 Nano base model checkpoint, training with this data improved a proxy SimpleQA score from 40.2 to 50.2.
- Nemotron-Pretraining-Moral-Scenarios: In the SFT data we previously released, we included multiple-choice questions about moral scenarios. These questions were constructed using situations and norms from Moral Stories and actions from Social Chemistry. We sampled a subset of these examples and created a chain-of-thought version using Qwen3-235B-A22B-Thinking-2507.
- Nemotron-Pretraining-Generative: Diverse large-scale synthetic questions with generative answers.
- Nemotron-Pretraining-Multiple-Choice: Diverse large-scale synthetic questions and answers in multiple choice format.
For more details about how the Nemotron-Pretraining-Generative and Nemotron-Pretraining-Multiple-Choice subsets were created, see Task-Seeded Synthetic Q&A Generation for Nemotron Pretraining.
The table below shows the number of tokens and the model used to generate these subsets:
The columns are as follows:
- text: The primary data field, containing the content to be used for pretraining.
- license: The license(s) governing the sample (e.g., ‘cc-by-4.0’).
- metadata: A dictionary detailing the following:
- category: Data type ('Nemotron-Pretraining-Generative' or 'Nemotron-Pretraining-Multiple-Choice').
- models_used: Models used to generate the data (e.g., '').
- uuid: The unique identifier for this dataset entry.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
05/18/2026
Version:
Nemotron-Pretraining-Specialized-v1.2
Previous Version(s):
Relationship to Previous Version(s): This is an extension of the previously released datasets, containing new data. They are meant to be used together.
License/Terms of Use:
The Nemotron-Pretraining-Fact-Seeking dataset is licensed under the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
The Nemotron-Pretraining-Moral-Scenarios dataset is licensed under the Creative Commons Attribution 4.0 International License (CC-BY-4.0). Additional Information: MIT License.
For the Nemotron-Pretraining-Multiple-Choice and Nemotron-Pretraining-Generative datasets, please see the 'license' entry for each sample. The majority of the datasets is licensed under the Creative Commons Attribution 4.0 International License (CC-BY-4.0). Some samples are licensed under the Creative Common Attribution 2.0 Generic License (CC-BY-2.0).
Each user is responsible for checking the content of datasets and the applicable licenses and determining if suitable for the intended use.
This dataset contains synthetic data created using the following models:
DeepSeek-v3, Mixtral-8x22B-v0.1, Qwen3-30B-A3B-Instruct-2507, Qwen3-235B-A22B, Qwen3-235B-A22B-Thinking-2507.
If the Nemotron-Pretraining-Multiple-Choice or Nemotron-Pretraining-Generative datasets are used to create, train, fine-tune, or otherwise improve an AI model, which is distributed or made available, such AI model may be subject to redistribution and use requirements in the DeepSeek License Agreement.
Intended Usage:
The Nemotron-Pre-Training-Specialized-v1.2 Dataset is intended to be used by the community to continue to improve open models.
Dataset Characterization
Data Collection Method
- Synthetic: Synthetic generation using large language models (DeepSeek-v3, Mixtral-8x22B-v0.1, Qwen3-30B-A3B-Instruct-2507, Qwen3-235B-A22B, Qwen3-235B-A22B-Thinking-2507).
Labeling Method
- Not Applicable
Dataset Format
Modality: Text
Format: Parquet
Dataset Quantification
Record Count: 599.5M samples
Measurement of Total Data Storage: 53.6 GB
Reference(s):
If you use our dataset in your research, please cite our NVIDIA Nemotron 3 Ultra tech report.
For more details on Nemotron-Pretraining-Generative and Nemotron-Pretraining-Multiple-Choice, please see Task-Seeded Synthetic Q&A Generation for Nemotron Pretraining.
Ethical Considerations:
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here.
