CoolFace
Datasetpublic

nvidia/Nemotron-RL-Lightning-Training-Blend

Dataset Description: This dataset provides the training-data blend used for the Reinforcement Learning with Verifiable Rewards (RLVR) stage of the public Nemotron-3.5-Lightning post-training recipe. The blend is consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. See the recipe for how the blend is used. The blend mixes NVIDIA-released datasets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Lightning-Training-Blend.

sourceHugging Facecc-by-4.0updated 29d agoView on Hugging Face
3likes521downloads

nvidia/Nemotron-RL-Lightning-Training-Blend · main · files are served by the source, never re-hosted here