CoolFace
20 results

rl-training

CopyleftCultivars /Agriculture-Agent-RL-Training-Data Agriculture Agent RL Training Data A growing dataset of RL rollout trajectories for LLM agents on natural/regenerative farming — the first RL/trajectory-shaped dataset in the Copyleft Cultivars collection (every prior dataset here is SFT/conversational Q&A). Agents call real tools (primarily cultivars-mcp, a plant-genomics MCP server) across 9 knowledge categories (plus a 10th, organic_chemistry_soil_science, added 2026-08-11, and an 11th, organic_chemistry_synthesis, added… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/Agriculture-Agent-RL-Training-Data.text-generation1 likes6.2k downloads25d agoHugging Facenvidia /Nemotron-RL-Ultra-Training-Blends Dataset Description: This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used. The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.tabulartext-generation10K<n<100K19 likes1.4k downloads2mo agoHugging Facenvidia /Nemotron-3-Nano-RL-Training-Blend Dataset Description: Nemotron-3-Nano-RL-Training-Blend is a curated dataset blend used to train the Nemotron-3-Nano-30B-A3B model. The blend consists of the following component datasets, with mixing ratios shown in parentheses: nvidia/Nemotron-RL-instruction_following (0.17) nvidia/Nemotron-RL-knowledge-mcqa (0.20) nvidia/Nemotron-RL-agent-workplace_assistant (0.10) nvidia/Nemotron-RL-instruction_following-structured_outputs (0.05) nvidia/Nemotron-RL-coding-competitive_coding… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-3-Nano-RL-Training-Blend.29 likes733 downloads9mo agoHugging Facenvidia /Nemotron-RL-Super-Training-Blends Dataset Description: Nemotron-3-Super-RL-Training-Blends contains the dataset blends used to train the Nemotron-3-Super-120B-A12B model. RL training for the Nemotron-3-Super-120B-A12B model is done in 6 stages: RLVR 1, RLVR 2, RLVR 3, SWE 1, SWE 2, and RLHF. The blends for each stage consist of data from various datasets, which we detail below. The percentages in parentheses indicate the mixing ratios of the dataset components. Note that the model was also trained on additional data… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Super-Training-Blends.39 likes681 downloads6mo agoHugging Facenvidia /Nemotron-RL-Lightning-Training-Blend Dataset Description: This dataset provides the training-data blend used for the Reinforcement Learning with Verifiable Rewards (RLVR) stage of the public Nemotron-3.5-Lightning post-training recipe. The blend is consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. See the recipe for how the blend is used. The blend mixes NVIDIA-released datasets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Lightning-Training-Blend.text-generation3 likes386 downloads25d agoHugging FaceTencentBAC /HyLaR_RL_Training_Dataset HyLaR RL Training Dataset This repository contains the reinforcement learning (RL) training dataset for HyLaR (Hybrid Latent Reasoning), as presented in the paper HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization. Resources Paper: HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization GitHub Repository: EthenCheng/HyLaR Model Checkpoint: HyLaR-Qwen2.5-VL-7B Dataset Description This dataset is designed for training… See the full description on the dataset page: https://huggingface.co/datasets/TencentBAC/HyLaR_RL_Training_Dataset.textimage-text-to-text10K<n<100K1 likes323 downloads3mo agoHugging Face