datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LIBERO_LeRobot_v3
LIBERO LeRobot v3
Dataset Summary
nvidia/LIBERO_LeRobot_v3 is a LeRobotDataset v3.0 conversion of the LIBERO robot manipulation benchmark. LIBERO is designed for studying lifelong robot learning and knowledge transfer across language-conditioned manipulation tasks. This repository packages the LIBERO task suites as LeRobot-compatible datasets with Parquet state/action data, MP4 video observations, and LeRobot metadata.
The dataset is organized as five top-level… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LIBERO_LeRobot_v3.PhysicalAI-Robotics-Manipulation-Kitchen
PhysicalAI Robotics Manipulation in the Kitchen
Dataset Description:
PhysicalAI-Robotics-Manipulation-Kitchen is a dataset of automatic generated motions of robots performing operations such as opening and closing cabinets, drawers, dishwashers and fridges. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Kitchen.miracl-vision
MIRACL-VISION
MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark.
This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.PhysicalAI-Robotics-GR00T-Teleop-G1
Unitree G1 Fruits Pick and Place 1K Dataset
Dataset Description:
The PhysicalAI-Robotics-GR00T-Teleop-G1 dataset consists of1000 teleoperation trajectories of real robot data using Unitree G1, with upper body control. The robot chooses the correct fruit to pick and place on the plate according to the language prompt. A total of 4 fruits are used: Apple, Pear, Starfruit, Grape. The robot is equipped with the default realsense camera, and a pair of Unitree G1 Tri-fingers… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-G1.Arena-G1-Loco-Manipulation-Task
Dataset Description:
The Arena-G1-Loco-Manipulation-Task dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (G1) loco-manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for box pick and place task.
Dataset Name
# Trajectories
G1 Loco-Manipulation Task
50
This dataset is ideal for behavior cloning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-G1-Loco-Manipulation-Task.GR00T-N1.7-AppleToPlate
Dataset Description:
The GR00T-N1.7-AppleToPlate dataset is a multimodal collection of trajectories collected on a Unitree G1 humanoid robot. It supports a humanoid (G1) static loco-manipulation task in which the robot picks up an apple and places it on a plate. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for an apple pick-and-place task.
Dataset Name
# Trajectories
G1 Static… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/GR00T-N1.7-AppleToPlate.aisimulate-fpm-dataset
AISimulate FPM Dataset
Forward Pass Model (FPM) libraries and independent latency measurements. This
private repository is the input-data source for FPM Gym.
Layout
data/
<org>--<model>/<system>/<framework>/<framework-version>/<parallelism>/
manifest.json # flat metadata for the promoted snapshot
fpm/ # optional self-benchmark FPM libraries and sidecars
provenance/ # FPM producer configurations… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/aisimulate-fpm-dataset.PhysicalAI-Robotics-Manipulation-ObjectsPhysicalAI-Robotics-Manipulation-Objects is a dataset of automatic generated motions of robots performing operations such as picking and placing objects in a kitchen environment. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with Kinova Gen3 arms. The environments are kitchen scenes where the furniture and appliances were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Objects.PhysicalAI-GR00T-Tuned-Tasks
Dataset Description:
This dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (GR1) tabletop manipulation tasks for industrial settings. Each dataset entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for tasks like pouring nuts or sorting pipes by color.
Dataset Name
# Trajectories
Exhaust-Pipe-Sorting-task
1000
Nut-Pouring-task
1000
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-GR00T-Tuned-Tasks.Arena-G1-Static-PickNPlace-Task
Dataset Description:
The Arena-G1-Static-PickNPlace-Task dataset is a multimodal collection of trajectories generated in Isaac Lab. It supports humanoid (G1) loco-manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for an apple pick-and-place task.
Dataset Name
# Trajectories
G1 Static PickNPlace Task
200
This dataset is ideal for behavior… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-G1-Static-PickNPlace-Task.Arena-GR1-Manipulation-PlaceItemCloseDoor-Task
Dataset Description:
The Arena-GR1-Manipulation-PlaceItemCloseDoor-Task dataset is a multimodal collection of trajectories generated in Isaac Lab. It supports humanoid (GR1) manipulation tasks in the IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, and action) needed to train and evaluate generalist robot policies for a sequential task (e.g. putting object into a fridge and closing the door).
Dataset Name
# Trajectories
GR1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-PlaceItemCloseDoor-Task.NV-Raw2Insights-US
NV-Raw2Insights-US Simulations
Dataset Description
NV-Raw2Insights-US Simulations is a simulated full synthetic aperture (FSA) ultrasound dataset for training and evaluating neural networks on sound speed estimation, phase aberration correction, and tissue segmentation.
Each sample is a single-frame FSA acquisition from a 180-element linear array simulated over a heterogeneous tissue phantom containing cysts. The dataset provides raw baseband IQ channel data alongside… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/NV-Raw2Insights-US.Arena-GR1-Manipulation-Task
Dataset Description:
The Arena-GR1-Manipulation-Task dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (GR1) manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for opening microwave task.
Dataset Name
# Trajectories
GR1 Manipulation Task
50
This dataset is ideal for behavior cloning, policy learning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-Task.NVIDIA_WORKSPACE_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 41,
"total_frames": 30151,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:41"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tremmelnicholas/NVIDIA_WORKSPACE_3.Arena-GR1-Manipulation-Task-v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "GR1",
"total_episodes": 50,
"total_frames": 4928,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-Task-v3.nvidia-math-vectorizedNVIDIA_WORKSPACEThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 27,
"total_frames": 19913,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:27"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tremmelnicholas/NVIDIA_WORKSPACE.Nemotron-RL-math-advanced_calculations
Dataset Description:
The Nemotron-RL-math-advanced_calculations is a dataset designed to test a model's ability to solve complex, multi-step math problems in a multi-step agentic environment. It involves counterintuitive calculations with varying levels of function composition.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-math-advanced_calculations.nvidialab_recogidaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/BravoRobots/nvidialab_recogida.nvidialab_recogida2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/BravoRobots/nvidialab_recogida2.nvidialab_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/BravoRobots/nvidialab_v1.vipe-web360
ViPE Dataset Release
This dataset contains the camera pose, depth, and intrinsics estimated using ViPE. For more details of the dataset, please refer to the Github link.
Please consider citing the following paper if you found this dataset helpful:
@article{huang2025vipe,
title={Vipe: Video pose engine for 3d geometric perception},
author={Huang, Jiahui and Zhou, Qunjie and Rabeti, Hesam and Korovko, Aleksandr and Ling, Huan and Ren, Xuanchi and Shen, Tianchang and Gao, Jun and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/vipe-web360.nvidialab_act_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/BravoRobots/nvidialab_act_v1.nvidialab_recogida4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/BravoRobots/nvidialab_recogida4.nvidialab_4000k_bgnegro_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/BravoRobots/nvidialab_4000k_bgnegro_v1.Nemotron-Cascade-RL-RLHF
Dataset Description:
The Nemotron-Cascade-RL-RLHF dataset is designed for Reinforcement Learning from Human Feedback (RLHF) training. It contains prompts and associated metadata to support the development of language model alignment.
This dataset is ready for commercial use.
The dataset contains the following subset:
RLHF Training Data
This data contains 45,882 samples used for RLHF training. It includes prompts, data sources, and category information.
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-RL-RLHF.nvidialab_recogida3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/BravoRobots/nvidialab_recogida3.NVIDIA_WORKSPACE_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 19,
"total_frames": 11155,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:19"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tremmelnicholas/NVIDIA_WORKSPACE_2.NVIDIA_WORKSPACE_MERGEDThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 56,
"total_frames": 37915,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:56"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tremmelnicholas/NVIDIA_WORKSPACE_MERGED.details_nvidia__Llama3-ChatQA-1.5-8B
Dataset Card for Evaluation run of nvidia/Llama3-ChatQA-1.5-8B
Dataset automatically created during the evaluation run of model nvidia/Llama3-ChatQA-1.5-8B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_nvidia__Llama3-ChatQA-1.5-8B.
