rl-training
Agriculture-Agent-RL-Training-Data
Agriculture Agent RL Training Data
A growing dataset of RL rollout trajectories for LLM agents on
natural/regenerative farming — the first RL/trajectory-shaped dataset in the
Copyleft Cultivars collection
(every prior dataset here is SFT/conversational Q&A). Agents call real tools
(primarily cultivars-mcp,
a plant-genomics MCP server) across 9 knowledge categories (plus a 10th,
organic_chemistry_soil_science, added 2026-08-11, and an 11th,
organic_chemistry_synthesis, added… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/Agriculture-Agent-RL-Training-Data.Nemotron-RL-Ultra-Training-Blends
Dataset Description:
This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used.
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.Nemotron-3-Nano-RL-Training-Blend
Dataset Description:
Nemotron-3-Nano-RL-Training-Blend is a curated dataset blend used to train the Nemotron-3-Nano-30B-A3B model. The blend consists of the following component datasets, with mixing ratios shown in parentheses:
nvidia/Nemotron-RL-instruction_following (0.17)
nvidia/Nemotron-RL-knowledge-mcqa (0.20)
nvidia/Nemotron-RL-agent-workplace_assistant (0.10)
nvidia/Nemotron-RL-instruction_following-structured_outputs (0.05)
nvidia/Nemotron-RL-coding-competitive_coding… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-3-Nano-RL-Training-Blend.Nemotron-RL-Super-Training-Blends
Dataset Description:
Nemotron-3-Super-RL-Training-Blends contains the dataset blends used to train the Nemotron-3-Super-120B-A12B model. RL training for the Nemotron-3-Super-120B-A12B model is done in 6 stages: RLVR 1, RLVR 2, RLVR 3, SWE 1, SWE 2, and RLHF. The blends for each stage consist of data from various datasets, which we detail below. The percentages in parentheses indicate the mixing ratios of the dataset components. Note that the model was also trained on additional data… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Super-Training-Blends.Nemotron-RL-Lightning-Training-Blend
Dataset Description:
This dataset provides the training-data blend used for the Reinforcement Learning with Verifiable Rewards (RLVR) stage of the public Nemotron-3.5-Lightning post-training recipe. The blend is consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. See the recipe for how the blend is used.
The blend mixes NVIDIA-released datasets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Lightning-Training-Blend.HyLaR_RL_Training_Dataset
HyLaR RL Training Dataset
This repository contains the reinforcement learning (RL) training dataset for HyLaR (Hybrid Latent Reasoning), as presented in the paper HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization.
Resources
Paper: HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization
GitHub Repository: EthenCheng/HyLaR
Model Checkpoint: HyLaR-Qwen2.5-VL-7B
Dataset Description
This dataset is designed for training… See the full description on the dataset page: https://huggingface.co/datasets/TencentBAC/HyLaR_RL_Training_Dataset.
