CoolFace
Datasetpublic

RoboCOIN/AI2_Alphabot_2_stack_cups

AI2_Alphabot_2_stack_cups Dataset Description This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Task Preview View Video Directly Overview Total Episodes: 403 Total Frames: 324466 FPS: 30 Dataset Size: 9.35 GB Robot Name: AI2_Alphabot_2 End-Effector Type: two_finger_end_effector Teleoperation Type: vr_controller Sensors: cam_front_chest_rgb, cam_front_head_rgb, cam_left_wrist_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_stack_cups.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes1.5kdownloads
Dataset Card

AI2Alphabot2stackcups

Dataset Description

This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.

Task Preview

<video src="videos/chunk-000/observation.images.camfrontchestrgb/episode000000.mp4" controls width="640"></video>

View Video Directly

Overview

  • Total Episodes: 403
  • Total Frames: 324466
  • FPS: 30
  • Dataset Size: 9.35 GB
  • Robot Name: AI2_Alphabot_2
  • End-Effector Type: two_finger_end_effector
  • Teleoperation Type: vr_controller
  • Sensors: cam_front_chest_rgb, cam_front_head_rgb, cam_left_wrist_rgb, cam_right_wrist_rgb
  • Camera Information: camfrontchestrgb; camfrontheadrgb; camleftwristrgb; camrightwristrgb
  • Scene: Catering->Restaurant
  • Objects: plastic cup(purple), plastic cup(blue), plastic cup(yellow)
  • Task Description: Stack the three plastic cups sequentially using one hand.

Primary Task Instruction

Stack the three plastic cups sequentially using one hand.

Robot Configuration

  • Robot Name: AI2_Alphabot_2
  • Codebase Version: v2.1
  • End-Effector Type: two_finger_end_effector
  • Teleoperation Type: vr_controller

Scene and Objects

Scene Type

Catering->Restaurant

Objects

  • plastic cup(purple)
  • plastic cup(blue)
  • plastic cup(yellow)

Task Descriptions

  • Standardized Task Description: Stack the three plastic cups sequentially using one hand.
  • Operation Type: fixed_single_arm
  • Environment Type: real_world

Sub-Tasks

This dataset includes 1 distinct subtasks:

  1. 1.Stack the three plastic cups sequentially using one hand.

Atomic Actions

  • pick
  • place

Hardware and Sensors

Sensors

  • cam_front_chest_rgb
  • cam_front_head_rgb
  • cam_left_wrist_rgb
  • cam_right_wrist_rgb

Camera Information

  • cam_front_chest_rgb: dtype=video, shape=480x640x3, resolution=640x480, codec=av1, pix_fmt=yuv420p
  • cam_front_head_rgb: dtype=video, shape=480x640x3, resolution=640x480, codec=av1, pix_fmt=yuv420p
  • cam_left_wrist_rgb: dtype=video, shape=480x640x3, resolution=640x480, codec=av1, pix_fmt=yuv420p
  • cam_right_wrist_rgb: dtype=video, shape=480x640x3, resolution=640x480, codec=av1, pix_fmt=yuv420p

Coordinate System

  • Definition: right-hand-frame

Dimensions & Units

  • Joint Rotation: radian
  • End-Effector Rotation: radian
  • End-Effector Translation: meter

Dataset Statistics

MetricValue
Total Episodes403
Total Frames324466
Total Tasks1
Total Videos1612
Total Chunks1
Chunk Size1000
FPS30
State Dimensions34
Action Dimensions34
Camera Views4
Dataset Size9.35 GB

Data Splits

The dataset is organized into the following splits:

  • Training: Episodes 0:402

Dataset Structure

This dataset follows the LeRobot format and contains the following components:

Data Files

  • Videos: Compressed video files containing RGB camera observations
  • State Data: Robot joint positions, velocities, and other state information
  • Action Data: Robot action commands and trajectories
  • Metadata: Episode metadata, timestamps, and annotations

File Organization

  • Data Path Pattern: data/chunk-{id}/episode_{id}.parquet
  • Video Path Pattern: videos/chunk-{id}/observation.images.cam_front_chest_rgb/episode_{id}.mp{id}
  • Chunking: Data is organized into 1 chunk(s) of size 1000

Data Structure (Tree)

AI2_Alphabot_2_stack_cups_0_qced_hardlink/
|-- annotations
|   |-- eef_acc_mag_annotation.jsonl
|   |-- eef_direction_annotation.jsonl
|   |-- eef_velocity_annotation.jsonl
|   |-- gripper_activity_annotation.jsonl
|   `-- gripper_mode_annotation.jsonl
|-- data
|   `-- chunk-000
|       |-- episode_000000.parquet
|       |-- episode_000001.parquet
|       |-- episode_000002.parquet
|       |-- episode_000003.parquet
|       |-- episode_000004.parquet
|       |-- episode_000005.parquet
|       |-- episode_000006.parquet
|       |-- episode_000007.parquet
|       |-- episode_000008.parquet
|       |-- episode_000009.parquet
|       |-- episode_000010.parquet
|       `-- episode_000011.parquet
|       `-- ... (391 more entries)
|-- meta
|   |-- episodes.jsonl
|   |-- episodes_stats.jsonl
|   |-- info.json
|   `-- tasks.jsonl
`-- videos
    `-- chunk-000
        |-- observation.images.cam_front_chest_rgb
        |-- observation.images.cam_front_head_rgb
        |-- observation.images.cam_left_wrist_rgb
        `-- observation.images.cam_right_wrist_rgb

Camera Views

This dataset includes 4 camera views: cam_front_chest_rgb, cam_front_head_rgb, cam_left_wrist_rgb, cam_right_wrist_rgb.

Features (Full YAML)

yaml
observation.images.cam_front_chest_rgb:
  dtype: video
  shape:
  - 480
  - 640
  - 3
  names:
  - height
  - width
  - channels
  info:
    video.height: 480
    video.width: 640
    video.codec: av1
    video.pix_fmt: yuv420p
    video.is_depth_map: false
    video.fps: 30
    video.channels: 3
    has_audio: false
observation.images.cam_front_head_rgb:
  dtype: video
  shape:
  - 480
  - 640
  - 3
  names:
  - height
  - width
  - channels
  info:
    video.height: 480
    video.width: 640
    video.codec: av1
    video.pix_fmt: yuv420p
    video.is_depth_map: false
    video.fps: 30
    video.channels: 3
    has_audio: false
observation.images.cam_left_wrist_rgb:
  dtype: video
  shape:
  - 480
  - 640
  - 3
  names:
  - height
  - width
  - channels
  info:
    video.height: 480
    video.width: 640
    video.codec: av1
    video.pix_fmt: yuv420p
    video.is_depth_map: false
    video.fps: 30
    video.channels: 3
    has_audio: false
observation.images.cam_right_wrist_rgb:
  dtype: video
  shape:
  - 480
  - 640
  - 3
  names:
  - height
  - width
  - channels
  info:
    video.height: 480
    video.width: 640
    video.codec: av1
    video.pix_fmt: yuv420p
    video.is_depth_map: false
    video.fps: 30
    video.channels: 3
    has_audio: false
observation.state:
  dtype: float32
  shape:
  - 34
  names:
  - left_arm_joint_1_rad
  - left_arm_joint_2_rad
  - left_arm_joint_3_rad
  - left_arm_joint_4_rad
  - left_arm_joint_5_rad
  - left_arm_joint_6_rad
  - left_arm_joint_7_rad
  - left_eef_pos_x_m
  - left_eef_pos_y_m
  - left_eef_pos_z_m
  - left_eef_rot_euler_x_rad
  - left_eef_rot_euler_y_rad
  - left_eef_rot_euler_z_rad
  - right_arm_joint_1_rad
  - right_arm_joint_2_rad
  - right_arm_joint_3_rad
  - right_arm_joint_4_rad
  - right_arm_joint_5_rad
  - right_arm_joint_6_rad
  - right_arm_joint_7_rad
  - right_eef_pos_x_m
  - right_eef_pos_y_m
  - right_eef_pos_z_m
  - right_eef_rot_euler_x_rad
  - right_eef_rot_euler_y_rad
  - right_eef_rot_euler_z_rad
  - left_gripper_open
  - right_gripper_open
  - neck_joint_1_rad
  - neck_joint_2_rad
  - torso_joint_1_rad
  - torso_joint_2_rad
  - torso_joint_3_rad
  - torso_joint_4_rad
action:
  dtype: float32
  shape:
  - 34
  names:
  - left_arm_joint_1_rad
  - left_arm_joint_2_rad
  - left_arm_joint_3_rad
  - left_arm_joint_4_rad
  - left_arm_joint_5_rad
  - left_arm_joint_6_rad
  - left_arm_joint_7_rad
  - left_eef_pos_x_m
  - left_eef_pos_y_m
  - left_eef_pos_z_m
  - left_eef_rot_euler_x_rad
  - left_eef_rot_euler_y_rad
  - left_eef_rot_euler_z_rad
  - right_arm_joint_1_rad
  - right_arm_joint_2_rad
  - right_arm_joint_3_rad
  - right_arm_joint_4_rad
  - right_arm_joint_5_rad
  - right_arm_joint_6_rad
  - right_arm_joint_7_rad
  - right_eef_pos_x_m
  - right_eef_pos_y_m
  - right_eef_pos_z_m
  - right_eef_rot_euler_x_rad
  - right_eef_rot_euler_y_rad
  - right_eef_rot_euler_z_rad
  - left_gripper_open
  - right_gripper_open
  - neck_joint_1_rad
  - neck_joint_2_rad
  - torso_joint_1_rad
  - torso_joint_2_rad
  - torso_joint_3_rad
  - torso_joint_4_rad
timestamp:
  dtype: float32
  shape:
  - 1
  names: null
frame_index:
  dtype: int64
  shape:
  - 1
  names: null
episode_index:
  dtype: int64
  shape:
  - 1
  names: null
index:
  dtype: int64
  shape:
  - 1
  names: null
task_index:
  dtype: int64
  shape:
  - 1
  names: null
gripper_open_scale_state:
  names:
  - left_gripper_open_scale
  - right_gripper_open_scale
  dtype: float64
  shape:
  - 2
gripper_open_scale_action:
  names:
  - left_gripper_open_scale
  - right_gripper_open_scale
  dtype: float64
  shape:
  - 2
eef_sim_pose_state:
  names:
  - left_eef_pos_x
  - left_eef_pos_y
  - left_eef_pos_z
  - left_eef_rot_x
  - left_eef_rot_y
  - left_eef_rot_z
  - right_eef_pos_x
  - right_eef_pos_y
  - right_eef_pos_z
  - right_eef_rot_x
  - right_eef_rot_y
  - right_eef_rot_z
  dtype: float32
  shape:
  - 12
eef_sim_pose_action:
  names:
  - left_eef_pos_x
  - left_eef_pos_y
  - left_eef_pos_z
  - left_eef_rot_x
  - left_eef_rot_y
  - left_eef_rot_z
  - right_eef_pos_x
  - right_eef_pos_y
  - right_eef_pos_z
  - right_eef_rot_x
  - right_eef_rot_y
  - right_eef_rot_z
  dtype: float32
  shape:
  - 12
eef_direction_state:
  names:
  - left_eef_direction
  - right_eef_direction
  dtype: int32
  shape:
  - 2
eef_direction_action:
  names:
  - left_eef_direction
  - right_eef_direction
  dtype: int32
  shape:
  - 2
eef_velocity_state:
  names:
  - left_eef_velocity
  - right_eef_velocity
  dtype: int32
  shape:
  - 2
eef_velocity_action:
  names:
  - left_eef_velocity
  - right_eef_velocity
  dtype: int32
  shape:
  - 2
eef_acc_mag_state:
  names:
  - left_eef_acc_mag
  - right_eef_acc_mag
  dtype: int32
  shape:
  - 2
eef_acc_mag_action:
  names:
  - left_eef_acc_mag
  - right_eef_acc_mag
  dtype: int32
  shape:
  - 2
gripper_mode_state:
  names:
  - left_gripper_mode
  - right_gripper_mode
  dtype: int32
  shape:
  - 2
gripper_mode_action:
  names:
  - left_gripper_mode
  - right_gripper_mode
  dtype: int32
  shape:
  - 2
gripper_activity_state:
  names:
  - left_gripper_activity
  - right_gripper_activity
  dtype: int32
  shape:
  - 2
gripper_activity_action:
  names:
  - left_gripper_activity
  - right_gripper_activity
  dtype: int32
  shape:
  - 2

Available Annotations

This dataset includes rich annotations to support diverse learning approaches:

  • eef_acc_mag_annotation.jsonl
  • eef_direction_annotation.jsonl
  • eef_velocity_annotation.jsonl
  • gripper_activity_annotation.jsonl
  • gripper_mode_annotation.jsonl

Dataset Tags

  • RoboCOIN
  • LeRobot

Authors

Contributors

This dataset is contributed by:-RoboCOIN Team at Beijing Academy of Artificial Intelligence (BAAI)

Annotators

No annotator information available.

Links

Contact and Support

For questions, issues, or feedback regarding this dataset, please contact us.

Support

For technical support, please open an issue on our GitHub repository.

License

apache-2.0

Citation

If you use this dataset in your research, please cite:

bibtex
@article{robocoin,
  title={RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation},
  author={Shihan Wu, Xuecheng Liu, Shaoxuan Xie, Pengwei Wang, Xinghang Li, Bowen Yang, Zhe Li, Kai Zhu, Hongyu Wu, Yiheng Liu, Zhaoye Long, Yue Wang, Chong Liu, Dihan Wang, Ziqiang Ni, Xiang Yang, You Liu, Ruoxuan Feng, Runtian Xu, Lei Zhang, Denghang Huang, Chenghao Jin, Anlan Yin, Xinlong Wang, Zhenguo Sun, Junkai Zhao, Mengfei Du, Mingyu Cao, Xiansheng Chen, Hongyang Cheng, Xiaojie Zhang, Yankai Fu, Ning Chen, Cheng Chi, Sixiang Chen, Huaihai Lyu, Xiaoshuai Hao, Yequan Wang, Bo Lei, Dong Liu, Xi Yang, Yance Jiao, Tengfei Pan, Yunyan Zhang, Songjing Wang, Ziqian Zhang, Xu Liu, Ji Zhang, Caowei Meng, Zhizheng Zhang, Jiyang Gao, Song Wang, Xiaokun Leng, Zhiqiang Xie, Zhenzhen Zhou, Peng Huang, Wu Yang, Yandong Guo, Yichao Zhu, Suibing Zheng, Hao Cheng, Xinmin Ding, Yang Yue, Huanqian Wang, Chi Chen, Jingrui Pang, YuXi Qian, Haoran Geng, Lianli Gao, Haiyuan Li, Bin Fang, Gao Huang, Yaodong Yang, Hao Dong, He Wang, Hang Zhao, Yadong Mu, Di Hu, Hao Zhao, Tiejun Huang, Shanghang Zhang, Yonghua Lin, Zhongyuan Wang and Guocai Yao},
  journal={arXiv preprint arXiv:2511.17441},
  url = {https://arxiv.org/abs/2511.17441},
  year={2025},
  }

Additional References

If you use this dataset, please also consider citing: LeRobot Framework: https://github.com/huggingface/lerobot

Version Information

Initial Release