datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HelpSteer2
HelpSteer2: Open-source dataset for training top-performing reward models
HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
This dataset has been created in partnership with Scale AI.
When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.PhysicalAI-Robotics-GR00T-Teleop-Sim
Simulation GR1 Tabletop Task 1K Dataset
Dataset Description:
The PhysicalAI-Robotics-GR00T-Teleop-GR1 dataset consists of 1000 teleoperation trajectories in simulation using the GR1 robot with upper body control. The simulation setup mimics tabletop manipulation tasks and uses RGB observations with a virtual camera. The robot is equipped with simulated Fourier hands.
This dataset is ready for non-commercial use.
Dataset Owner(s):
NVIDIA GEAR… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-Sim.PhysicalAI-Robotics-GR00T-Teleop-GR1
Introduction
TL;DR: DreamDojo is a generalist robot world model pretrained on 44k hours of human egocentric data, showing unprecedented generalization to diverse objects and environments.
Project page: https://dreamdojo-world.github.io/
Paper: https://arxiv.org/abs/2602.06949
Code: https://github.com/NVIDIA/DreamDojo
How to Use
Check out https://github.com/NVIDIA/DreamDojo
Citation
@article{gao2026dreamdojo,
title={DreamDojo: A Generalist Robot… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.Nemotron-ClimbMix
ClimbMix Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.Granary
Granary: Speech Recognition and Translation Dataset in 25 European Languages
Granary is a large-scale, open-source multilingual speech dataset covering 25 European languages for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) tasks.
Overview
Granary addresses the scarcity of high-quality speech data for low-resource languages by consolidating multiple datasets under a unified framework:
🗣️ ~1M hours of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Granary.HelpSteer
HelpSteer: Helpfulness SteerLM Dataset
HelpSteer is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
Leveraging this dataset and SteerLM, we train a Llama 2 70B to reach 7.54 on MT Bench, the highest among models trained on open-source datasets based on MT Bench Leaderboard as of 15 Nov 2023.
This model is available on… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer.PhysicalAI-Robotics-Manipulation-Kitchen
PhysicalAI Robotics Manipulation in the Kitchen
Dataset Description:
PhysicalAI-Robotics-Manipulation-Kitchen is a dataset of automatic generated motions of robots performing operations such as opening and closing cabinets, drawers, dishwashers and fridges. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Kitchen.LIBERO_LeRobot_v3
LIBERO LeRobot v3
Dataset Summary
nvidia/LIBERO_LeRobot_v3 is a LeRobotDataset v3.0 conversion of the LIBERO robot manipulation benchmark. LIBERO is designed for studying lifelong robot learning and knowledge transfer across language-conditioned manipulation tasks. This repository packages the LIBERO task suites as LeRobot-compatible datasets with Parquet state/action data, MP4 video observations, and LeRobot metadata.
The dataset is organized as five top-level… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LIBERO_LeRobot_v3.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838 different… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.Nemotron-RL-Ultra-Training-Blends
Dataset Description:
This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used.
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.miracl-vision
MIRACL-VISION
MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark.
This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.PhysicalAI-Robotics-GR00T-Teleop-G1
Unitree G1 Fruits Pick and Place 1K Dataset
Dataset Description:
The PhysicalAI-Robotics-GR00T-Teleop-G1 dataset consists of1000 teleoperation trajectories of real robot data using Unitree G1, with upper body control. The robot chooses the correct fruit to pick and place on the plate according to the language prompt. A total of 4 fruits are used: Apple, Pear, Starfruit, Grape. The robot is equipped with the default realsense camera, and a pair of Unitree G1 Tri-fingers… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-G1.aerial-isac-srs-iq
Aerial ISAC SRS I/Q
Raw uplink Sounding Reference Signal (SRS) I/Q captured on the
NVIDIA Aerial
5G testbed, paired with a synchronized video and camera-derived ground truth for
two pedestrians and a car moving through the sensing area.
The labeled span in real time: camera view, Range-Doppler map, and range/velocity
tracks. Also available as isac_rd_demo.mp4.
Dataset Description:
This dataset provides synchronized multi-modal recordings designed for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/aerial-isac-srs-iq.Nemotron-RL-Agentic-SWE-Pivot-v1
Dataset Description:
The SWE-RL dataset provides GitHub issues for training and validating real-world software engineering agents using the OpenHands environment in NeMo Gym. The dataset is a refactored version of the SWE-Gym and R2E-Gym datasets to support the NeMo Gym input format.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.Arena-G1-Loco-Manipulation-Task
Dataset Description:
The Arena-G1-Loco-Manipulation-Task dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (G1) loco-manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for box pick and place task.
Dataset Name
# Trajectories
G1 Loco-Manipulation Task
50
This dataset is ideal for behavior cloning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-G1-Loco-Manipulation-Task.hifitts-2
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
Dataset Description
This repository contains the metadata for HiFiTTS-2, a large scale speech dataset derived from LibriVox audiobooks. For more details, please refer to our paper.
The dataset contains metadata for approximately 36.7k hours of audio from 5k speakers that can be downloaded from LibriVox at a 48 kHz sampling rate.
The metadata contains estimated bandwidth, which can be used to infer the original… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/hifitts-2.GR00T-N1.7-AppleToPlate
Dataset Description:
The GR00T-N1.7-AppleToPlate dataset is a multimodal collection of trajectories collected on a Unitree G1 humanoid robot. It supports a humanoid (G1) static loco-manipulation task in which the robot picks up an apple and places it on a plate. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for an apple pick-and-place task.
Dataset Name
# Trajectories
G1 Static… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/GR00T-N1.7-AppleToPlate.Nemotron-Cascade-2-RL-data
Dataset Description:
The Nemotron-Cascade-2-RL dataset is a curated reinforcement learning (RL) dataset blend used to train Nemotron-Cascade-2-30B-A3B model. It includes instruction-following RL, multi-domain RL, on-policy distillation, and software engineering RL (SWE-RL) data.
This dataset is ready for commercial use.
The dataset contains the following subset:
IF-RL
Contains 45,879 training samples for instruction-following RL. Our curation process mainly… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-RL-data.compute-eval
Dataset Card for ComputeEval
ComputeEval is a benchmark for evaluating LLM-generated CUDA code on correctness and performance. Each problem provides a self-contained programming challenge — spanning kernels, runtime APIs, memory management, parallel algorithms, GPU libraries, and template/DSL kernel authoring — with a held-out test harness for functional validation and optional performance benchmarks for measuring GPU execution time against a baseline.
Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/compute-eval.PhysicalAI-GR00T-Tuned-Tasks
Dataset Description:
This dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (GR1) tabletop manipulation tasks for industrial settings. Each dataset entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for tasks like pouring nuts or sorting pipes by color.
Dataset Name
# Trajectories
Exhaust-Pipe-Sorting-task
1000
Nut-Pouring-task
1000
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-GR00T-Tuned-Tasks.PhysicalAI-Robotics-Manipulation-ObjectsPhysicalAI-Robotics-Manipulation-Objects is a dataset of automatic generated motions of robots performing operations such as picking and placing objects in a kitchen environment. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with Kinova Gen3 arms. The environments are kitchen scenes where the furniture and appliances were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Objects.Arena-G1-Static-PickNPlace-Task
Dataset Description:
The Arena-G1-Static-PickNPlace-Task dataset is a multimodal collection of trajectories generated in Isaac Lab. It supports humanoid (G1) loco-manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for an apple pick-and-place task.
Dataset Name
# Trajectories
G1 Static PickNPlace Task
200
This dataset is ideal for behavior… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-G1-Static-PickNPlace-Task.aisimulate-fpm-dataset
AISimulate FPM Dataset
Forward Pass Model (FPM) libraries and independent latency measurements. This
private repository is the input-data source for FPM Gym.
Layout
data/
<org>--<model>/<system>/<framework>/<framework-version>/<parallelism>/
manifest.json # flat metadata for the promoted snapshot
fpm/ # optional self-benchmark FPM libraries and sidecars
provenance/ # FPM producer configurations… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/aisimulate-fpm-dataset.Nemotron-RL-Instruction-Following-MultiTurnChat-v1
Dataset Description:
The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.aerial-isac-pusch-hest
Aerial ISAC PUSCH Channel Estimates
Uplink PUSCH DMRS channel estimates from a 5G CBRS cell running indoors on the
NVIDIA Aerial
testbed, paired with camera-derived floor positions of a person walking through the
cell: 1.19 million estimates over 18 runs, nine with a person in the area and nine
recorded empty as a background reference.
Plan view in the label coordinate frame, with the recorded track of a clear run
and of an obstacle run.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/aerial-isac-pusch-hest.NV-Raw2Insights-US
NV-Raw2Insights-US Simulations
Dataset Description
NV-Raw2Insights-US Simulations is a simulated full synthetic aperture (FSA) ultrasound dataset for training and evaluating neural networks on sound speed estimation, phase aberration correction, and tissue segmentation.
Each sample is a single-frame FSA acquisition from a 180-element linear array simulated over a heterogeneous tissue phantom containing cysts. The dataset provides raw baseband IQ channel data alongside… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/NV-Raw2Insights-US.Arena-GR1-Manipulation-PlaceItemCloseDoor-Task
Dataset Description:
The Arena-GR1-Manipulation-PlaceItemCloseDoor-Task dataset is a multimodal collection of trajectories generated in Isaac Lab. It supports humanoid (GR1) manipulation tasks in the IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, and action) needed to train and evaluate generalist robot policies for a sequential task (e.g. putting object into a fridge and closing the door).
Dataset Name
# Trajectories
GR1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-PlaceItemCloseDoor-Task.Arena-GR1-Manipulation-Task
Dataset Description:
The Arena-GR1-Manipulation-Task dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (GR1) manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for opening microwave task.
Dataset Name
# Trajectories
GR1 Manipulation Task
50
This dataset is ideal for behavior cloning, policy learning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-Task.glm52-fidelity-nvfp4-nvidia-v1
fidelity--glm52.malaiwah.quant.nvfp4-nvidia
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from nvidia/GLM-5.2-NVFP4.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-nvfp4-nvidia-v1.Arena-DROID-Camera-Sensitivity-Workflow-Sample
Arena DROID Camera Sensitivity Workflow Sample
Dataset Description
Arena-DROID-Camera-Sensitivity-Workflow-Sample is a compact set of episode-level results generated by an Isaac Lab-Arena simulation experiment. It lets users run the documented camera sensitivity analysis without first executing the policy-evaluation sweep.
The experiment evaluates an OpenPI pi05 policy on a DROID Rubik's-cube pick-and-place task while independently varying the wrist-camera… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-DROID-Camera-Sensitivity-Workflow-Sample.
