datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Cascade-2-SFT-Data
Nemotron-Cascade-2-SFT-Data
We release the SFT data used for training Nemotron-Cascade-2.
Data sources
Math
Our non-proof math prompts are sourced from Nemotron-Cascade-1-SFT and Nemotron-Math-v2, with responses generated by DeepSeek-V3.2, DeepSeek-V3.2-Speciale, and GPT-OSS-120B. For mathematical proofs, prompts are taken from Nemotron-Math-Proofs-v1 and generated using DeepSeek-V3.2-Speciale.
Science
We collect science prompts from… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-SFT-Data.Nemotron-Cascade-SFT-Stage-2
Nemotron-Cascade-SFT-Stage-2
Supervised fine-tuning (SFT) for Nemotron-Cascade is performed in two stages. The Stage-1 SFT focuses on the math, code, science, and general domains, leveraging a broad and diverse collection of data sources. The Stage-2 SFT further expands coverage to include math, code, science, tool calling, software engineering (SWE), instruction following, and general domains.
In Stage-2, the math domain leverages questions from OpenMathReasoning. The code domain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-SFT-Stage-2.Nemotron-Cascade-2-RL-data
Dataset Description:
The Nemotron-Cascade-2-RL dataset is a curated reinforcement learning (RL) dataset blend used to train Nemotron-Cascade-2-30B-A3B model. It includes instruction-following RL, multi-domain RL, on-policy distillation, and software engineering RL (SWE-RL) data.
This dataset is ready for commercial use.
The dataset contains the following subset:
IF-RL
Contains 45,879 training samples for instruction-following RL. Our curation process mainly… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-RL-data.Nemotron-Cascade-2-SFT-Data
Nemotron-Cascade-2-SFT-Data
We release the SFT data used for training Nemotron-Cascade-2.
Data sources
Math
Our non-proof math prompts are sourced from Nemotron-Cascade-1-SFT and Nemotron-Math-v2, with responses generated by DeepSeek-V3.2, DeepSeek-V3.2-Speciale, and GPT-OSS-120B. For mathematical proofs, prompts are taken from Nemotron-Math-Proofs-v1 and generated using DeepSeek-V3.2-Speciale.
Science
We collect science prompts from… See the full description on the dataset page: https://huggingface.co/datasets/BrunoN-Dev/Nemotron-Cascade-2-SFT-Data.Nemotron-Cascade-2-SFT-Data
Nemotron-Cascade-2-SFT-Data
We release the SFT data used for training Nemotron-Cascade-2.
Data sources
Math
Our non-proof math prompts are sourced from Nemotron-Cascade-1-SFT and Nemotron-Math-v2, with responses generated by DeepSeek-V3.2, DeepSeek-V3.2-Speciale, and GPT-OSS-120B. For mathematical proofs, prompts are taken from Nemotron-Math-Proofs-v1 and generated using DeepSeek-V3.2-Speciale.
Science
We collect science prompts from… See the full description on the dataset page: https://huggingface.co/datasets/febcheema/Nemotron-Cascade-2-SFT-Data.Nemotron-Cascade-2-SFT-Data-Small
Nemotron-Cascade-2-SFT-Data-Small
A 20% random sample of nvidia/Nemotron-Cascade-2-SFT-Data, merged into a single train split with 4,898,804 rows.
Subsets included (all merged)
Original subset
Files sampled
~Rows sampled
math
math_notool, math_proof, math_tool
~1,045,266
science
science
~544,383
chat
chat_part_1 – chat_part_4
~2,794,866
instruction_following
instruction_following
~163,869
safety
safety
~693
conversational_agent
conversational_agent… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Nemotron-Cascade-2-SFT-Data-Small.nemotron-cascade2-cheating-attempts
Nemotron-Cascade-2 30B A3B — Cheating Investigation
Tool the model had: a single tool, stateful_python_code_exec
(Jupyter sandbox). Network egress from inside that sandbox is not blocked
— the model can urllib.request.urlopen, requests.get, and even
pip install from PyPI.
Baseline accuracy across the 10,197 traces: 5,748 / 10,197 = 56.4 %.
1. Headline numbers
Bucket
Traces
Correct
Accuracy
Δ vs baseline
All traces
10,197
5,748
56.4 %
—
Any network… See the full description on the dataset page: https://huggingface.co/datasets/chankhavu/nemotron-cascade2-cheating-attempts.Nemotron-Cascade-2-SFT-Data-9k-subsetNemotron-Cascade-2-SFT-Data-16k-subsetnemotron-cascade-2-science-dedupednemotron-cascade-2-science-deduped-200knemotron-cascade-2-science-deduped-100k
