datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-ClimbMix
ClimbMix Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.fineinstructions_nemotron
✨ Note: For all FineInstructions resources please visit: https://huggingface.co/fineinstructions
This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline.
The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details.
Each .parquet file in the data folder has a corresponding judge-*.json file that… See the full description on the dataset page: https://huggingface.co/datasets/fineinstructions/fineinstructions_nemotron.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838 different… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.Nemotron-RL-Ultra-Training-Blends
Dataset Description:
This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used.
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.Nemotron-RL-Agentic-SWE-Pivot-v1
Dataset Description:
The SWE-RL dataset provides GitHub issues for training and validating real-world software engineering agents using the OpenHands environment in NeMo Gym. The dataset is a refactored version of the SWE-Gym and R2E-Gym datasets to support the NeMo Gym input format.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.Nemotron-Research-Reasoning-Qwen-1.5B_eval_569aNemotron-Cascade-2-RL-data
Dataset Description:
The Nemotron-Cascade-2-RL dataset is a curated reinforcement learning (RL) dataset blend used to train Nemotron-Cascade-2-30B-A3B model. It includes instruction-following RL, multi-domain RL, on-policy distillation, and software engineering RL (SWE-RL) data.
This dataset is ready for commercial use.
The dataset contains the following subset:
IF-RL
Contains 45,879 training samples for instruction-following RL. Our curation process mainly… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-RL-data.OpenReasoning-Nemotron-7B_eval_8179
mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
79.0
98.8
89.0
81.7
60.1
62.5
50.6
46.8
68.7
13.3
49.6
59.7
AIME24
Average Accuracy: 79.00% ± 1.42%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
70.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179.AceReason-Nemotron-7B_eval_118b
mlfoundations-dev/AceReason-Nemotron-7B_eval_118b
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_official
Average Accuracy: 43.85% ± 0.24%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
43.37%
121
279
2
44.09%
123
279
3
44.09%
123
279
Nemotron-SFT-Agentic-v2-prompt-only
Nemotron-SFT-Agentic-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.Nemotron-RL-Instruction-Following-MultiTurnChat-v1
Dataset Description:
The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.nemotron-mc-en-ar-midtrain
nemotron-mc-en-ar-midtrain
Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.nemotron-r1-en-ar-midtrain
nemotron-r1-en-ar-midtrain
Arabic translation of the Llama_Nemotron_Post_Training_Dataset_reasoning_r1 split of smoltalk2 (config Mid, pinned revision fc6cc21): reasoning traces with <think> blocks in a conversational format. Translated with RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic (greedy) on H100s. FP8 was verified lossless against its bf16 parent before the run (chrF 96.4, 0 of 510 chunks materially diverged). All 3,644,790 source rows are present, none dropped. Sibling… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-r1-en-ar-midtrain.nemotron_cc_v2_hq_packed4096_200shard
Nemotron-CC-v2 High-Quality, packed to 4096 tokens (train)
Documents from nvidia/Nemotron-CC-v2
High-Quality subset, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS
appended per document) and greedily packed into sequences of at most 4096 tokens.
A document is never split across a pack boundary; documents longer than 4096
are truncated to their own pack. Every pack ends on an EOS/document boundary.
Schema
index (int64): running pack id… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096_200shard.OpenReasoning-Nemotron-1.5B_eval_8179
mlfoundations-dev/OpenReasoning-Nemotron-1.5B_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
49.7
83.0
78.0
49.4
31.0
35.5
19.8
14.6
40.7
12.0
24.3
32.3
AIME24
Average Accuracy: 49.67% ± 1.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
50.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-1.5B_eval_8179.nemotron_cc_v2_hq_packed4096
Nemotron-CC-v2 High-Quality, packed to 4096 tokens
5% subset of nvidia/Nemotron-CC-v2
High-Quality documents, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS
appended per document) and greedily packed into sequences of at most 4096 tokens.
A document is never split across a pack boundary; documents longer than 4096
are truncated to their own pack. Every pack ends on an EOS/document boundary.
Schema
index (int64): running pack id
input_ids… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096.Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.exp-pool-nemotron-math-dolma2-tokenized
Locus EXP Nemotron Math - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.AceReason-Nemotron-7B_eval_c64a
mlfoundations-dev/AceReason-Nemotron-7B_eval_c64a
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_v3
Average Accuracy: 41.67% ± 0.45%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
41.04%
110
268
2
41.42%
111
268
3
42.54%
114
268
nemotron-sft-balanced-2b-v1
Nemotron SFT Dataset
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Statistics
Total Samples: 200,000
Total Tokens: 1,252,287,904
Average Tokens per Sample: 6261.4
Tokenizer: Qwen/Qwen3-0.6B
Random Seed: 42
Strategy: balanced
Subset Distribution
Subset
Samples
Tokens
Target
Completion
Avg Tokens/Sample
Stage-1/math
20,000
151,546,125
20,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-2b-v1.Nemotron-RLHF-GenRM-v1
Dataset Description:
This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking.
The dataset is composed of:
Preference data focused on diverse domains
A synthetic safety blend
The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.Ornith-1.5-35B-A3B-Nemotron-v2-100M
Ornith 1.5 35B A3B Nemotron v2 100M
This dataset contains 108,729 English conversations with
108,729 regenerated assistant turns and 100,014,884 generated
assistant completion tokens. 100M refers to the completion-token target, not the
number of examples.
The prompt mix is a deterministic sample from
nvidia/Nemotron-Post-Training-Dataset-v2.
It covers the source dataset's chat, code, math, and STEM subsets. Every assistant turn
was regenerated with ornith-ai/Ornith-1.5-35B-A3B;… See the full description on the dataset page: https://huggingface.co/datasets/jzinno/Ornith-1.5-35B-A3B-Nemotron-v2-100M.NVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scoresgenerations-nemotron-nano-9b-v2-simnpo-gentle-igm-10bnemotron-nano-eval-logs-and-scoresgenerations-nemotron-nano-9b-v2-simnpo-gentle-baselinenemotron-sft-general-focused-stage1-2-ChatML-V3
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 496,385
Total Tokens: 1,114,218,401
Average Tokens per Sample: 2244.7
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V3.pretrain-nemotron-math-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
22,927,812,461 (22.9B)
Trainable tokens
22,927,812,461 (22.9B)
Documents
21,377,358
Shards
180
UTF-8 bytes
77,994,866,327
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix.Nemotron-Reason2-GPTgenerations-nemotron-nano-9b-v2-simnpo-baseline
