CoolFace
Datasetpublic

nvidia/Nemotron-RL-Lightning-Training-Blend

Dataset Description: This dataset provides the training-data blend used for the Reinforcement Learning with Verifiable Rewards (RLVR) stage of the public Nemotron-3.5-Lightning post-training recipe. The blend is consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. See the recipe for how the blend is used. The blend mixes NVIDIA-released datasets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Lightning-Training-Blend.

sourceHugging Facecc-by-4.0updated 29d agoView on Hugging Face
3likes521downloads
6 commits on main
262eb5829d ago

Update README.md

leannachr
c6616041mo ago

Update README.md

yfw
e9591631mo ago

Update README.md

yfw
7ed9ba12mo ago

Correct recipe URL to nemotron-3.5-lightning.md

yfw
01cd1d42mo ago

Add files using upload-large-folder tool

yfw
47715732mo ago

initial commit

yfw