cascade
Datasets
All datasets matching “cascade”cascade-eval-pool
cascade eval pool — lagged public reveal (exact bytes)
Retired snapshots of the held-out evaluation pool used by the
cascade subnet. Each folder is a
byte-identical mirror of the pool/snapshots/block-<N>.tar that validators
scored — downloaded from the private pool bucket, sha256-verified against the
publisher index, and republished unmodified. A snapshot is revealed only after
a newer snapshot has superseded it, so no revealed pool can be selected by a
current or future round.… See the full description on the dataset page: https://huggingface.co/datasets/Tensor-Link/cascade-eval-pool.Nemotron-Cascade-2-SFT-Data
Nemotron-Cascade-2-SFT-Data
We release the SFT data used for training Nemotron-Cascade-2.
Data sources
Math
Our non-proof math prompts are sourced from Nemotron-Cascade-1-SFT and Nemotron-Math-v2, with responses generated by DeepSeek-V3.2, DeepSeek-V3.2-Speciale, and GPT-OSS-120B. For mathematical proofs, prompts are taken from Nemotron-Math-Proofs-v1 and generated using DeepSeek-V3.2-Speciale.
Science
We collect science prompts from… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-SFT-Data.cascade-testnet-mirrorNemotron-Cascade-SFT-Stage-2
Nemotron-Cascade-SFT-Stage-2
Supervised fine-tuning (SFT) for Nemotron-Cascade is performed in two stages. The Stage-1 SFT focuses on the math, code, science, and general domains, leveraging a broad and diverse collection of data sources. The Stage-2 SFT further expands coverage to include math, code, science, tool calling, software engineering (SWE), instruction following, and general domains.
In Stage-2, the math domain leverages questions from OpenMathReasoning. The code domain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-SFT-Stage-2.Nemotron-Cascade-2-RL-data
Dataset Description:
The Nemotron-Cascade-2-RL dataset is a curated reinforcement learning (RL) dataset blend used to train Nemotron-Cascade-2-30B-A3B model. It includes instruction-following RL, multi-domain RL, on-policy distillation, and software engineering RL (SWE-RL) data.
This dataset is ready for commercial use.
The dataset contains the following subset:
IF-RL
Contains 45,879 training samples for instruction-following RL. Our curation process mainly… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-RL-data.fire-fusion-cascades-500m
FireFusion Cascades 500m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over the Eastern Cascades of Washington State, a 272 km square running from the Cascade crest through the Okanogan Highlands, the most fire-active terrain in the state. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 500m by 500m grid covering every fire season 2003-2020.
Daily fire-season… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-cascades-500m.
