CoolFace
Datasetpublic

mickeykang/dvla-can-250hz-events-250to25fps-delta-trim

This dataset was created using LeRobot. Dataset Description dvla-can-250hz-events (250 -> 25 fps) - event frames accumulating 40 ms each 1016 episodes / 135,828 frames. Same scenes as the 250 Hz RGB set, rendered at 250 Hz, passed through the v2e DVS emulator (1.5.1, thresholds 0.15/0.15, slow-motion interpolation disabled because the frames are genuinely 250 Hz), then downsampled 10x: each output frame carries the events from a 40 ms window as a polarity image.… See the full description on the dataset page: https://huggingface.co/datasets/mickeykang/dvla-can-250hz-events-250to25fps-delta-trim.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes703downloads
Dataset Card

This dataset was created using LeRobot.

Dataset Description

dvla-can-250hz-events (250 -> 25 fps) - event frames accumulating 40 ms each

1016 episodes / 135,828 frames. Same scenes as the 250 Hz RGB set, rendered at 250 Hz, passed through the v2e DVS emulator (1.5.1, thresholds 0.15/0.15, slow-motion interpolation disabled because the frames are genuinely 250 Hz), then downsampled 10x: each output frame carries the events from a 40 ms window as a polarity image.

Two extra video columns on top of the 3 RGB cameras: observation.images.opst_cam_events, observation.images.wrist_cam_events.

The point of this variant: decide at 25 Hz - where the control-frequency penalty documented in the repo does not apply - while still seeing the 250 Hz motion between decisions. Median event coverage is ~2.5% of the frame for opst_cam and ~4.1% for wrist_cam.

Event frames are polarity renderings, not natural images, so ImageNet normalization must be turned off for those columns.

Task

MuJoCo / robosuite (Panda + OSC-or-IK) place: pick a can off the table and drop it into a bowl. In 80% of the episodes the can is rolling when the episode starts (speed sampled from 0.25-0.75 m/s); the rest are static. Demonstrations come from a scripted state machine that reads privileged simulator state, and only successful attempts are kept.

Layout

LeRobot v2.1, robot_type: panda, 3 RGB cameras at 360x480 (wrist_cam, side_cam, opst_cam). Training uses opst_cam + wrist_cam only.

  • —action (10,) = [dx, dy, dz, sin/cos of the 3 euler angles, gripper]
  • —observation.state (9,) = [x, y, z, sin/cos of the 3 euler angles]
  • —observation.environment_state (9,) = privileged object pose/velocity. Dropped during training (it is not available on a real robot).

action is a delta: the absolute target minus the state of that same frame. At evaluation the policy output is decoded as target = live_state + predicted_delta, re-anchored every step.

The absolute target it encodes is a future-EE relabel: instead of the pose the state machine commanded at time t, the label is the pose the arm had actually reached 320 ms later. A sweep over that offset found it matters enormously - replaying the labels open-loop succeeds 20/20 at 320 ms but 5/20 at 200 ms and 0/20 at 40 ms.

-delta-trim additionally drops the leading frames where the arm is still holding position, so that action[0] is a real motion rather than a placeholder (which otherwise causes a cold-start stall).

meta/camera.jsonl maps each episode_index back to its source HDF5 filename. Conversion shards round-robin, so episode order is not source order and this file is the only way back.

In-domain eval scenes (eval_scenes/)

Reproducing the in-domain numbers needs the scenes restored exactly as they were generated, which run_ckpt_eval*.sh does through --replay_h5_dir on the source HDF5s. Those are 37 GB (25 Hz) / 267 GB (250 Hz) and are not published.

They do not have to be. eval_server.load_h5_episode reads only the root attrs (env_config, seed, instruction, init_*) and the action dataset -- the camera frames, which are 99.99% of the file, are re-rendered by the server. Stripped to that, a 250 Hz 20-scene set goes from 5.1 GB to 0.37 MB, and ships here under eval_scenes/. Restoring from the stripped files was verified to reproduce the original scene exactly: object pose and velocity, bowl pose, arm pose and action length all match.

bash
huggingface-cli download mickeykang/dvla-can-250hz-events-250to25fps-delta-trim --repo-type dataset --include 'eval_scenes/*' --local-dir .
REPLAY_DIR=$PWD/eval_scenes/indomain_scenes_s20_250 ... bash overnight_delta_n3/run_policy_eval.sh

Out-of-distribution scenes ship here too, under eval_scenes/ood_scenes_s20_250. They are normally sampled from seeds 9000000-9000019 at run time, but that only reproduces on the machine that produced the published numbers: the sampler's bowl-clearance test reads mesh radii out of assets/, and the result also depends on the robosuite/MuJoCo build and the floating-point path. Replaying the frozen scenes removes all of that. The dumped scene names match the ones in the original eval logs, and replaying reproduces the seed-sampled scene exactly (object, bowl and joints identical; end-effector differs by 3e-08 because init_* is stored float32, a property of the original pipeline).

bash
REPLAY_DIR=$PWD/eval_scenes/ood_scenes_s20_250 ... bash overnight_delta_n3/run_policy_eval.sh

Reproducing

Generation, conversion, training and evaluation scripts, plus the measured results and the gotchas that cost the most time, are documented at https://github.com/mickeykang16/DynamicVLA/tree/mujoco.

  • —Homepage: https://github.com/mickeykang16/DynamicVLA/tree/mujoco
  • —Paper: [More Information Needed]
  • —License: apache-2.0

Dataset Structure

meta/info.json:

json
{
    "codebase_version": "v2.1",
    "robot_type": "panda",
    "total_episodes": 1016,
    "total_frames": 135828,
    "total_tasks": 3,
    "total_videos": 5080,
    "total_chunks": 2,
    "chunks_size": 1000,
    "fps": 25,
    "splits": {
        "train": "0:1016"
    },
    "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
    "video_path": "videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4",
    "features": {
        "action": {
            "dtype": "float32",
            "shape": [
                10
            ],
            "names": [
                "ee_pos_x",
                "ee_pos_y",
                "ee_pos_z",
                "ee_rot_x_sin",
                "ee_rot_x_cos",
                "ee_rot_y_sin",
                "ee_rot_y_cos",
                "ee_rot_z_sin",
                "ee_rot_z_cos",
                "gripper"
            ]
        },
        "observation.state": {
            "dtype": "float32",
            "shape": [
                9
            ],
            "names": [
                "ee_pos_x",
                "ee_pos_y",
                "ee_pos_z",
                "ee_rot_x_sin",
                "ee_rot_x_cos",
                "ee_rot_y_sin",
                "ee_rot_y_cos",
                "ee_rot_z_sin",
                "ee_rot_z_cos"
            ]
        },
        "observation.environment_state": {
            "dtype": "float32",
            "shape": [
                9
            ],
            "names": [
                "o0",
                "o1",
                "o2",
                "o3",
                "o4",
                "o5",
                "o6",
                "o7",
                "o8"
            ]
        },
        "observation.images.wrist_cam": {
            "dtype": "video",
            "shape": [
                360,
                480,
                3
            ],
            "names": [
                "height",
                "width",
                "channel"
            ],
            "info": {
                "video.height": 360,
                "video.width": 480,
                "video.codec": "h264",
                "video.pix_fmt": "yuv420p",
                "video.is_depth_map": false,
                "video.fps": 25,
                "video.channels": 3,
                "has_audio": false
            }
        },
        "observation.images.side_cam": {
            "dtype": "video",
            "shape": [
                360,
                480,
                3
            ],
            "names": [
                "height",
                "width",
                "channel"
            ],
            "info": {
                "video.height": 360,
                "video.width": 480,
                "video.codec": "h264",
                "video.pix_fmt": "yuv420p",
                "video.is_depth_map": false,
                "video.fps": 25,
                "video.channels": 3,
                "has_audio": false
            }
        },
        "observation.images.opst_cam": {
            "dtype": "video",
            "shape": [
                360,
                480,
                3
            ],
            "names": [
                "height",
                "width",
                "channel"
            ],
            "info": {
                "video.height": 360,
                "video.width": 480,
                "video.codec": "h264",
                "video.pix_fmt": "yuv420p",
                "video.is_depth_map": false,
                "video.fps": 25,
                "video.channels": 3,
                "has_audio": false
            }
        },
        "observation.images.wrist_cam_events": {
            "dtype": "video",
            "shape": [
                360,
                480,
                3
            ],
            "names": [
                "height",
                "width",
                "channel"
            ],
            "info": {
                "video.height": 360,
                "video.width": 480,
                "video.codec": "h264",
                "video.pix_fmt": "yuv420p",
                "video.is_depth_map": false,
                "video.fps": 25,
                "video.channels": 3,
                "has_audio": false
            }
        },
        "observation.images.opst_cam_events": {
            "dtype": "video",
            "shape": [
                360,
                480,
                3
            ],
            "names": [
                "height",
                "width",
                "channel"
            ],
            "info": {
                "video.height": 360,
                "video.width": 480,
                "video.codec": "h264",
                "video.pix_fmt": "yuv420p",
                "video.is_depth_map": false,
                "video.fps": 25,
                "video.channels": 3,
                "has_audio": false
            }
        },
        "timestamp": {
            "dtype": "float32",
            "shape": [
                1
            ],
            "names": null
        },
        "frame_index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        },
        "episode_index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        },
        "index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        },
        "task_index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        }
    }
}

Citation

BibTeX:

bibtex
[More Information Needed]