CoolFace
Datasetpublic

geonmin-kim/SO101-lv4-3color-cube-mat-to-mat-no-human-reset-v1-topcamcrop-fps15

This dataset was created using LeRobot. FPS 15 Even/Odd Split This is a derived version of geonmin-kim/SO101-lv4-3color-cube-mat-to-mat-no-human-reset-v1-topcamcrop (fps=30) resampled to fps=15 by splitting every original episode into two: original episode i (n frames) -> new episode 2*i = its even frames (0, 2, 4, ...; ceil(n/2) frames) original episode i -> new episode 2*i + 1 = its odd frames (1, 3, 5, ...; floor(n/2) frames) 246 original episodes become 492… See the full description on the dataset page: https://huggingface.co/datasets/geonmin-kim/SO101-lv4-3color-cube-mat-to-mat-no-human-reset-v1-topcamcrop-fps15.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes81downloads
Dataset Card

This dataset was created using LeRobot.

FPS 15 Even/Odd Split

This is a derived version of geonmin-kim/SO101-lv4-3color-cube-mat-to-mat-no-human-reset-v1-topcamcrop (fps=30) resampled to fps=15 by splitting every original episode into two:

  • —original episode i (n frames) -> new episode 2*i = its even frames (0, 2, 4, ...; ceil(n/2) frames)
  • —original episode i -> new episode 2*i + 1 = its odd frames (1, 3, 5, ...; floor(n/2) frames)

246 original episodes become 492 episodes; the total frame count (79,145) is unchanged. Timestamps in both the data parquet and the videos are renormalized to the fps=15 grid (frame_index / 15; the odd split's original 1/30 s offset is dropped). Videos were rebuilt by reordering decoded frames (even segment then odd segment per original episode, original per-camera file packing kept) and re-encoded at 15 fps with the original settings (AV1 / SVT-AV1, crf 30, g 2, preset 12, yuv420p). All per-episode and global stats (meta/stats.json, episodes parquet) were recomputed from the new data and frames.

target_cube_start_x/y are copied to both split episodes of an original episode (the cube does not move between original frames 0 and 1) and remain in cropped 592x384 top-camera pixels.

Top Camera Crop

This is a derived version of geonmin-kim/SO101-lv4-3color-cube-mat-to-mat-no-human-reset-v1 in which every observation.images.top frame is cropped to the workspace region (both black mats + robot, floor/background removed):

  • —crop rectangle in the original 640x480 frame: x=24, y=96, w=592, h=384
  • —resulting top-camera resolution: 592x384 (observation.images.wrist is unchanged, 640x480)
  • —videos re-encoded with the original settings (AV1 / SVT-AV1, crf 30, g 2, preset 12, yuv420p); frame counts and timestamps are identical to the source
  • —observation.images.top stats in meta/stats.json and per-episode stats in the episodes parquet were recomputed from the cropped frames
  • —target_cube_start_x / target_cube_start_y are expressed in the cropped frame (original values minus 24 / 96)

<a class="flex" href="https://huggingface.co/spaces/lerobot/visualize_dataset?path=geonmin-kim/SO101-lv4-3color-cube-mat-to-mat-no-human-reset-v1-topcamcrop-fps15"> <img class="block dark:hidden" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl.svg"/> <img class="hidden dark:block" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl-dark.svg"/> </a>

Dataset Description

  • —Homepage: [More Information Needed]
  • —Paper: [More Information Needed]
  • —License: apache-2.0

Dataset Structure

meta/info.json:

json
{
    "codebase_version": "v3.0",
    "fps": 15,
    "features": {
        "action": {
            "dtype": "float32",
            "shape": [
                6
            ],
            "names": [
                "shoulder_pan.pos",
                "shoulder_lift.pos",
                "elbow_flex.pos",
                "wrist_flex.pos",
                "wrist_roll.pos",
                "gripper.pos"
            ]
        },
        "observation.state": {
            "dtype": "float32",
            "shape": [
                6
            ],
            "names": [
                "shoulder_pan.pos",
                "shoulder_lift.pos",
                "elbow_flex.pos",
                "wrist_flex.pos",
                "wrist_roll.pos",
                "gripper.pos"
            ]
        },
        "observation.images.wrist": {
            "dtype": "video",
            "shape": [
                480,
                640,
                3
            ],
            "names": [
                "height",
                "width",
                "channels"
            ],
            "info": {
                "is_depth_map": false,
                "video.height": 480,
                "video.width": 640,
                "video.codec": "av1",
                "video.pix_fmt": "yuv420p",
                "video.fps": 15,
                "video.channels": 3,
                "has_audio": false,
                "video.g": 2,
                "video.crf": 30,
                "video.preset": 12,
                "video.fast_decode": 0,
                "video.video_backend": "pyav",
                "video.extra_options": {}
            }
        },
        "observation.images.top": {
            "dtype": "video",
            "shape": [
                384,
                592,
                3
            ],
            "names": [
                "height",
                "width",
                "channels"
            ],
            "info": {
                "is_depth_map": false,
                "video.height": 384,
                "video.width": 592,
                "video.codec": "av1",
                "video.pix_fmt": "yuv420p",
                "video.fps": 15,
                "video.channels": 3,
                "has_audio": false,
                "video.g": 2,
                "video.crf": 30,
                "video.preset": 12,
                "video.fast_decode": 0,
                "video.video_backend": "pyav",
                "video.extra_options": {}
            }
        },
        "timestamp": {
            "dtype": "float32",
            "shape": [
                1
            ],
            "names": null
        },
        "frame_index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        },
        "episode_index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        },
        "index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        },
        "task_index": {
            "dtype": "int64",
            "shape": [
                1
            ],
            "names": null
        }
    },
    "total_episodes": 492,
    "total_frames": 79145,
    "total_tasks": 6,
    "chunks_size": 1000,
    "data_files_size_in_mb": 100,
    "video_files_size_in_mb": 200,
    "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
    "video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
    "robot_type": "so_follower",
    "splits": {
        "train": "0:492"
    }
}

Target Cube Start Coordinates

meta/episodes/chunk-000/file-000.parquet contains three additional per-episode columns with the target cube's position in the first frame of each episode, measured in the cropped observation.images.top camera (592x384 pixel coordinates):

columndtypedescription
target_cube_colorstringTarget cube color parsed from the task (red / green / blue)
target_cube_start_xfloat32Cube centroid x in the episode's first cropped top-camera frame (pixels)
target_cube_start_yfloat32Cube centroid y in the episode's first cropped top-camera frame (pixels)

Coordinates were obtained (in the original 640x480 frames, then shifted by the crop offset) by color segmentation (per-color RGB ratio thresholds + compact connected-component filtering via OpenCV connectedComponentsWithStats, desk region y>=96 only) on the first top-camera frame of each episode, and all 246 detections were visually verified.

The exact script used to produce (and reproduce) these columns is included in the repo: `scripts/detect_target_cube_start_coords.py`.

bash
python scripts/detect_target_cube_start_coords.py --root /path/to/dataset --write-parquet

Citation

BibTeX:

bibtex
[More Information Needed]