cascade-rl
Nemotron-Cascade-2-RL-data
Dataset Description:
The Nemotron-Cascade-2-RL dataset is a curated reinforcement learning (RL) dataset blend used to train Nemotron-Cascade-2-30B-A3B model. It includes instruction-following RL, multi-domain RL, on-policy distillation, and software engineering RL (SWE-RL) data.
This dataset is ready for commercial use.
The dataset contains the following subset:
IF-RL
Contains 45,879 training samples for instruction-following RL. Our curation process mainly… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-RL-data.Nemotron-Cascade-RL-SWE
Dataset Description:
The Nemotron-Cascade-RL-SWE dataset is the RL training data for SWE code repairing task, consisting of SWE-Bench-Train, SWE-reBench, SWE-Smith, R2E-Gym/R2E-Gym-Subset and SWE-Fixer-Train.
We select the training data for SFT and RL stages based on its difficulty.
Also, to avoid data contamination, we exclude all instances originating from repositories present in the SWE-Bench_Verified evaluation dataset.
We create the prompts following the agentless mini… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-RL-SWE.Nemotron-Cascade-RL-Math
Nemotron-Cascade-RL-Math
Nemotron-Cascade-RL-Math is a diverse and high-quality dataset focused on math reasoning. It serves as the Math RL data for Nemotron-Cascade.
Nemotron-Cascade-RL-MATH contains 14,476 math problems and short answers, covering the data sources from OpenMathReasoning, NuminaMath-CoT, DeepScaleR, AceReason-Math. We conduct data decontamination and filter the sample that has a 9-gram overlap with any test sample in our math benchmarks.
The following are… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-RL-Math.Nemotron-Cascade-RL-RLHF
Dataset Description:
The Nemotron-Cascade-RL-RLHF dataset is designed for Reinforcement Learning from Human Feedback (RLHF) training. It contains prompts and associated metadata to support the development of language model alignment.
This dataset is ready for commercial use.
The dataset contains the following subset:
RLHF Training Data
This data contains 45,882 samples used for RLHF training. It includes prompts, data sources, and category information.
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-RL-RLHF.Nemotron-Cascade-2-RL-reproduction
Nemotron-Cascade-2 RL — Unified Reconstruction Recipe & Schema Sample
⚠️ これは NVIDIA 公式リリースではありません。 Nemotron-Cascade-2(arXiv:2603.19220)の
RL(事後学習)データを、公開済みの Nemotron 系データから再構成するための「レシピ+統一スキーマ」
パッケージです。同梱の train.jsonl は構造確認用の合成サンプルで、本物の学習データではありません
(各行 meta.synthetic_placeholder = true)。本物は build_cascade2_rl_data.py --mode full で
各 Nemotron データセットを取得して生成します。
Dataset Summary
Cascade 2 の RL は 7 段(IF-RL → Multi-domain RL → MOPD → RLHF → Long-context RL → Code RL →… See the full description on the dataset page: https://huggingface.co/datasets/TeamDelta/Nemotron-Cascade-2-RL-reproduction.RL_With_Cascade{'basic_science': 5000,
'coding_data': 3410,
'math_instruct': 500,
'creative_ideation': 4000,
'story_generation': 4000,
'creative_writing': 2098,
'summarization': 4000,
'gsm8k': 5000,
'general': 7500}
