datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgiBot-g1_left_capture_part
AgiBot-g1_left_capture_part
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: ruantong_a2d
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
factory
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
📊 Dataset Statistics
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_left_capture_part.Cobot_Magic_twist_bottle_cap
Cobot_Magic_twist_bottle_cap
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
twist
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_twist_bottle_cap.storageguertin-mcro-forensic-corpus-capture-sessions
Guertin MCRO Forensic Corpus: Capture Sessions
Contents: 16 OneWayVideo screen-capture sessions (started 2025-11-19 to 2026-05-13) — per session a full recording (video_full.mp4), WebDataset shards of hash-chained bundles (16 tars: screenshots, bundle JSON with network events and downloads, downloaded PDFs, OpenTimestamps proofs and time-stamp records), session reports (CSV/PDF) and event logs (data/).
Layout: <session>/video_full.mp4; <session>/shards/<session>-00000.tar… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-capture-sessions.video_captioningcaptain_cook_4degocentric-vr-capture-20h-multimodal-sample
Egocentric VR Capture — 20-Hour Multimodal Inspection Sample
195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.AgiBotWorld-Beta_G1_task_707_twist_the_bottle_cap
agibot_task_707
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 拧瓶盖
total_episodes: 1157
total_tasks: 1
size: 65G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_707_twist_the_bottle_cap.Lelouch_Vi_Britannia_FramePack_First_Last_Frame_Video_Captioned
AgiBot-g1_right_capture_part
AgiBot-g1_right_capture_part
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: ruantong_a2d
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
factory
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
📊 Dataset Statistics
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_right_capture_part.R4R-Auto-Eval
R4R Auto Eval
持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。
队友请先阅读 PROJECT_STATUS.md,然后按需查看:
benchmarks/:固定的视频输入、来源记录和分层标签;
pipelines/:判定方法及冻结配置;
runs/:不可覆盖的实验记录;
reports/:工作日志、方法分析和结果限制;
registry/:benchmark、pipeline 和 run 的机器可读索引。
当前范围
multiscene30 是 pipeline 开发集,不是干净的留出测试集;
reassemble40 是来自两个长录像的接触密集型校准集;
当前标签为来源数据提供方标签,尚未全部完成独立人工裁决;
Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测;
在完成逐来源许可证核查前,本仓库应保持 private。
当前发布版本:0.1.0。
AgiBotWorld-Beta_G1_task_431_Boil_coffee_with_a_capsule_machine
agibot_task_431
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 用胶囊机煮咖啡
total_episodes: 674
total_tasks: 1
size: 35G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_431_Boil_coffee_with_a_capsule_machine.Cobot_Magic_cap_the_pen_a
Cobot_Magic_cap_the_pen_a
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
insert
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_cap_the_pen_a.AgiBotWorld-Beta_G1_task_613_Insert_pen_cap
agibot_task_613
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 插入笔帽
total_episodes: 601
total_tasks: 1
size: 28G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_613_Insert_pen_cap.AIRBOT_MMK2_screw_the_bottle_cap
AIRBOT_MMK2_screw_the_bottle_cap
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_screw_the_bottle_cap.Interaction-Caption-Supplement
Interaction-Caption-Supplement
Dense, per-second caption data for training streaming video-language models — the kind that
must decide when to speak and when to stay silent while a video is still playing.
It is a supplement to momo321654/VL-Interaction-EN:
that release is strong on chat / event_grounding but its narration task is 98.6 %
ASR-derived (Live-WhisperX-526K + OmniStar-RNG), i.e. the supervision describes what a speaker
said, not what is visible. This supplement adds… See the full description on the dataset page: https://huggingface.co/datasets/momo321654/Interaction-Caption-Supplement.R1_Lite_move_the_position_of_the_coffee_capsule
R1_Lite_move_the_position_of_the_coffee_capsule
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
place
pick
grasp
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_move_the_position_of_the_coffee_capsule.put_coffee_cap_teaboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 40,
"total_frames": 14806,
"total_tasks": 1,
"total_videos": 80,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lirislab/put_coffee_cap_teabox.pick_coffee_capsule_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5",
"total_episodes": 770,
"total_frames": 452236,
"total_tasks": 84,
"total_videos": 1540,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:770"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/pick_coffee_capsule_merged.so101-red-cap-pick-place-demo-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/dharak22/so101-red-cap-pick-place-demo-v2.kbot_cappuccino_countThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "kbot_right_arm_follower",
"total_episodes": 224,
"total_frames": 166988,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:224"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/infoslack/kbot_cappuccino_count.OmniPhysics-Caption_benchmark
OmniFysics-Captioner: Grounding Omni-Modal Understanding in the Physical World for Better Captioning
🌐 Project •
📈 Benchmark Overview •
🧪 What OPC Measures •
📊 Daily-Physics 50K Subset •
📚 Citation
Introduction
Building omni-modal models with physical intelligence requires benchmarks that
test whether generated captions preserve information from both visual and
audio streams. Existing detailed-caption benchmarks provide strong visual… See the full description on the dataset page: https://huggingface.co/datasets/Fysics-AI/OmniPhysics-Caption_benchmark.video_caption_datasetsso101-red-cap-pick-place-exact-like-lerobot_20260826_110927This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/dharak22/so101-red-cap-pick-place-exact-like-lerobot_20260826_110927.capx_dream_demoput_caps_into_teaboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 10759,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lirislab/put_caps_into_teabox.egocentric-vr-capture-1h-multimodal-sample
Egocentric VR Capture — 1-Hour Multimodal Inspection Sample
13 real-world task episodes / 108,029 frames / approximately 60 minutes captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible dataset is an inspection slice produced by the EXYLOS real-world data pipeline. It demonstrates capture quality, synchronization, schema, and QA metadata… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/egocentric-vr-capture-1h-multimodal-sample.unscrew-the-bottle-cap-0821This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 100,
"total_frames": 84419,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Xense/unscrew-the-bottle-cap-0821.Kanon_Videos_Omni_Captioned_1
G1_Dex1_Open_Bottle_Cap
