rov
Datasets
All datasets matching “rov”RoVid-X
Rethinking Video Generation Model for the Embodied World
If you like our project, please give us a star ⭐ on GitHub for the latest update.
Key features
4M robotic video clips(10K+ hours) for large-scale video generation training.
1300+ fine-grained robotic skills, covering diverse actions and task primitives.
Multi-modal physical annotations, including RGB, depth, and optical flow.
Multi-robot and multi-task diversity… See the full description on the dataset page: https://huggingface.co/datasets/DAGroup-PKU/RoVid-X.ROVER
The ROVER Visual SLAM Benchmark
Paper
News
[2024/12/05] Initial code release.
[2025/05/20] ROVER is accepted to IEEE Transactions on Robotics.
[2025/05/25] Dataset released on HuggingFace, see Utility section for HuggingFace download script.
Getting Started
The only required software is Docker. Each SLAM method comes with its own Docker container, making setup straightforward. We recommend using VSCode with the Docker extension for an enhanced… See the full description on the dataset page: https://huggingface.co/datasets/iis-esslingen/ROVER.jointavbench
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
Overview
JointAVBench is a benchmark for evaluating omni-modal large language models on joint audio-visual reasoning tasks. Each multiple-choice question is designed to require both visual and auditory information.
This repository contains the audited release of JointAVBench under the roverx12345 namespace. The benchmark keeps the original 2,853-question split while refining answer… See the full description on the dataset page: https://huggingface.co/datasets/roverx12345/jointavbench.rovi
Dataset Card for ROVI
Dataset Summary
ROVI is the first language-driven video inpainting dataset. Please check our project page for more details.
To prevent data contamination and reduce evaluation time. We only provide part of the testing data in test.json.
@article{wu2024lgvi,
title={Towards language-driven video inpainting via multimodal large language models},
author={Wu, Jianzong and Li, Xiangtai and Si, Chenyang and Zhou, Shangchen and Yang, Jingkang and Zhang… See the full description on the dataset page: https://huggingface.co/datasets/jianzongwu/rovi.gemma_rover_scoop_up_to_5
gemma_rover_scoop_up_to_5
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
office-rover-2task-raw
office-rover-2task-raw
LeRobot v2.1 dataset for a single SO-101 arm: "pick up the red / green cube
and put it in the box" with both cubes on the table. Two language-conditioned
tasks (instructions in Russian). Built for fine-tuning NVIDIA Isaac GR00T N1.7
on one 24 GB RTX 4090 — recipe, patches and 1800 evaluated attempts:
https://github.com/VShirokun/gr00t-on-4090
What is in it: the raw, un-engineered demonstrations: cubes in fixed orientation, 240×320 cameras. Trained as-is… See the full description on the dataset page: https://huggingface.co/datasets/VShirokun/office-rover-2task-raw.
