mrtonyy/2cmCubeInBox_Grid
This dataset was created using LeRobot. Camera Setup Two USB UVC webcams, both captured at 640×480 @ 30fps: observation.images.overhead — Creative Live! Cam Sync 1080p V2. Mounted above the workspace, looking straight down at the grid, cube, and box. observation.images.wrist — Innomaker U20CAM-1080P-S1. Mounted on the follower arm's gripper, giving a close-up view of the grasp point that moves with the arm. Both cameras are 1080p-capable but recorded at 640×480 to match… See the full description on the dataset page: https://huggingface.co/datasets/mrtonyy/2cmCubeInBox_Grid.
This dataset was created using LeRobot.
<a class="flex" href="https://huggingface.co/spaces/lerobot/visualizedataset?path=mrtonyy/2cmCubeInBoxGrid"> <img class="block dark:hidden" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl.svg"/> <img class="hidden dark:block" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl-dark.svg"/> </a>
Dataset Description
- Homepage: [More Information Needed]
- Paper: [More Information Needed]
- License: apache-2.0
Camera Setup
Two USB UVC webcams, both captured at 640×480 @ 30fps:
- `observation.images.overhead` — Creative Live! Cam Sync 1080p V2. Mounted above the workspace, looking straight down at the grid, cube, and box.
- `observation.images.wrist` — Innomaker U20CAM-1080P-S1. Mounted on the follower arm's gripper, giving a close-up view of the grasp point that moves with the arm.
Both cameras are 1080p-capable but recorded at 640×480 to match the policy's input resolution.
Training Positions
The workspace was marked with a physical 2cm×2cm tape grid (matching the cube's own footprint), with one cell tagged as a fixed origin reference.
Cube-pick demonstrations cover 23 grid positions, ~10 repetitions each (230 episodes), forming a single dense, contiguous block directly above the destination box and below the arm's base, with no deliberately-skipped interior cells. Every cell inside that block was demonstrated; positions outside it were never shown to the policy during training. The remaining 30 episodes are "hold position, no cube present" idle demonstrations, teaching the policy to stay still when nothing is in view, for 260 episodes total.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6
]
},
"observation.state": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6
]
},
"observation.images.overhead": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"is_depth_map": false,
"video.height": 480,
"video.width": 640,
"video.codec": "av1",
"video.pix_fmt": "yuv420p",
"video.fps": 30,
"video.channels": 3,
"has_audio": false,
"video.g": 2,
"video.crf": 30,
"video.preset": 12,
"video.fast_decode": 0,
"video.video_backend": "pyav",
"video.extra_options": {}
}
},
"observation.images.wrist": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"is_depth_map": false,
"video.height": 480,
"video.width": 640,
"video.codec": "av1",
"video.pix_fmt": "yuv420p",
"video.fps": 30,
"video.channels": 3,
"has_audio": false,
"video.g": 2,
"video.crf": 30,
"video.preset": 12,
"video.fast_decode": 0,
"video.video_backend": "pyav",
"video.extra_options": {}
}
},
"timestamp": {
"dtype": "float32",
"shape": [
1
],
"names": null
},
"frame_index": {
"dtype": "int64",
"shape": [
1
],
"names": null
},
"episode_index": {
"dtype": "int64",
"shape": [
1
],
"names": null
},
"index": {
"dtype": "int64",
"shape": [
1
],
"names": null
},
"task_index": {
"dtype": "int64",
"shape": [
1
],
"names": null
}
},
"total_episodes": 260,
"total_frames": 141948,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"robot_type": "so_follower",
"splits": {
"train": "0:260"
}
}Citation
BibTeX:
[More Information Needed]