datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
so101-eval-gallerywat.ai-so101so101_tb4_pick_place_QIso101_tb4_RJ_testing_RB0so101_self_rollout_rawso101_isr_193
SO-101 Teleop — ISR-Standardized (193 episodes)
Training arm B of the ISR vs Partner-sampling SmolVLA experiment. This is the client's SO-101
teleop set with ISR (Information-Standardized Trajectory Resampling) applied — pauses removed,
motion kept at uniform information spacing — for comparison against the raw baseline and the
partner company's length-normalization method.
TL;DR
value
Source
makermods/200ep_blue_cube_orange_box (SO-101, real teleop)… See the full description on the dataset page: https://huggingface.co/datasets/Kavin60606/so101_isr_193.LeRobot-SO101-Pick-Place
LeRobot-SO101-Pick-Place
Dataset Summary
LeRobot dataset for SO101 pick-and-place with three sponge manipulation tasks and fixed/random object layouts. The dataset is intended for imitation learning, behavior cloning, policy evaluation, and robustness studies under layout changes.
Supported Tasks
Imitation learning
Behavior cloning
Policy evaluation
Robustness to layout changes
Dataset Structure
The dataset is organized into three… See the full description on the dataset page: https://huggingface.co/datasets/aswinkumar99/LeRobot-SO101-Pick-Place.wat.ai-so101-videosTPSoSe2026_Dataset_Collection_LeRobot_SO101
SO-101 Early Collection (superseded)
The project's first, miscellaneous recordings, made before the team settled on four
fixed tasks, a consistent recording protocol, and systematic prompt variation.
[!WARNING]
This dataset is superseded and not recommended for training. It is retained for
provenance and to document the project's history. The only model trained on it —
SmolVLA V1 Misc —
does not work.
Part of Project-IRA — Interactive Robotic Arm.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/Project-IRA/TPSoSe2026_Dataset_Collection_LeRobot_SO101.so101-teleop-vialsgso-so101-nexus
GSO objects for so101-nexus
Twelve models mirrored from Google's Scanned Objects
dataset
(Google Scanned Objects, "GSO"), CC-BY-4.0, via the real-world-scale OBJ +
PNG mirror at
kevinzakka/mujoco_scanned_objects
(mirror tooling MIT-licensed; the mesh/texture assets themselves stay
CC-BY-4.0 from Google).
Layout matches ai-habitat/ycb: one directory per model id under meshes/,
holding model.obj and texture.png.
Consumed by so101_nexus.gso_assets
(SO101_GSO_HF_REPO environment… See the full description on the dataset page: https://huggingface.co/datasets/johnsutor/gso-so101-nexus.AHA-WAM-SO101-HIL-training-assets
AHA-WAM SO101 Plug HIL Training Assets
Reproducibility bundle for the SO101 power-adapter insertion experiments.
It contains a compact, ZIP-based representation of the paths expected by the
AHA-WAM training configuration:
the released AHA-WAM-pretrained.pt initialization checkpoint;
so101_ahawam_plug.zip (the downloader extracts only task 02 and 04);
ahawam_hil_raw.zip (raw rich-v1/v2/v3 HG-DAgger recordings);
the fixed 02+04 action/state normalization statistics;
cached T5… See the full description on the dataset page: https://huggingface.co/datasets/Jill111/AHA-WAM-SO101-HIL-training-assets.so101_color_pen_sortso101_pick_place_batch2so101_camera_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"image": {
"dtype": "image",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"is_depth_map": false
}… See the full description on the dataset page: https://huggingface.co/datasets/codenmood/so101_camera_dataset.so101-smolvla-data
so101-smolvla-data — two SO-101 halves in LeRobot v3.0, ready for SmolVLA
std_mm_teleop_v30.tar 95 MB 100 eps 9,500 frames 1 task ISR-standardized real teleop
ego_v30.tar 319 MB 324 eps 36,442 frames 89 tasks retargeted egocentric video
std_results/ the standardization run that produced the first half
Both tars unpack to a complete LeRobot v3.0 tree (meta/ data/ videos/) that loads with
LeRobotDataset(repo_id, root=...). Same… See the full description on the dataset page: https://huggingface.co/datasets/angkul07/so101-smolvla-data.so101_sock_stowing2_boxesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 59,
"total_frames": 34574,
"total_tasks": 1,
"total_videos": 177,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:59"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/littledragon/so101_sock_stowing2_boxes.so101_tb4_dynamichandover_scene1_receiver_blue_brush_080726so101_car_pick_and_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 98,
"total_frames": 67279,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:98"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jonathm126/so101_car_pick_and_place.so101_pick_lift_cubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 1771,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gtgando/so101_pick_lift_cube.so101-teleop-vials-v2so101-segmentation
SO-101 Instance Segmentation Dataset
Instance segmentation dataset for the SO-101 robot arm, combining MuJoCo domain-randomised synthetic renders with hand-labelled real frames.
This dataset can be used to create a segmentation model for tracking the movements of the SO-101
Categories
8 classes, all prefixed so101_:
ID
Name
Notes
1
so101_base
2
so101_shoulder
3
so101_upper_arm
4
so101_lower_arm
5
so101_wrist
6
so101_gripper
includes moving… See the full description on the dataset page: https://huggingface.co/datasets/riversnow/so101-segmentation.so101_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 1791,
"total_tasks":1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yyiiii/so101_test.SO101-cap_stack_RGBblock_on_bluedish_10fps
SO101 CAP Stack RGB Blocks on Blue Dish
This dataset contains 100 LeRobot v3.0 demonstration episodes for an SO101 follower robot. The task is: Stack red, green, and blue blocks on the blue dish from bottom to top. The dataset was collected at 10 Hz and includes paired top-view and wrist-view RGB videos, robot state/action trajectories, and CAP skill annotations.
Dataset Details
Field
Value
Repository… See the full description on the dataset page: https://huggingface.co/datasets/CoRL2026-CSI/SO101-cap_stack_RGBblock_on_bluedish_10fps.MimicAnno-so101-26B-runs
SO101 annotations — Gemma-4 26B (QLoRA) MimicAnno run
20 episodes from the SO101 dataset, automatically segmented and
phase-labeled by the MimicAnno
pipeline using the 26B Gemma-4 QLoRA adapter
(https://huggingface.co/Gayagaya/gem4_26B_adapter).
Contents (per episode dir)
file
description
video.mp4
source episode video
signals.json
per-frame hand/object signals (schema v3)
boundaries.json
segment boundaries
annotation.jsonVLM phase labels + tools… See the full description on the dataset page: https://huggingface.co/datasets/Gayagaya/MimicAnno-so101-26B-runs.so101_camera_dataset_0016This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"image": {
"dtype": "image",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
]
},
"wrist_image": {
"dtype": "image"… See the full description on the dataset page: https://huggingface.co/datasets/codenmood/so101_camera_dataset_0016.so101_pick_place
Pick object and place in box
Project KIWI | Qian Group HRI Lab | University of Houston
Detail
Value
Task
Pick object and place in box
Episodes
50
FPS
30
Cameras
Gripper (Arducam) + Overhead (Intel RealSense)
Robot
SO-101 (Feetech STS3215)
Compute
NVIDIA Jetson Orin Nano Super
Framework
LeRobot
so101_barcode_ep_v5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/klvarshaa/so101_barcode_ep_v5.so101-robosta-cubemolmoact2-so101-zero-shot-eval
MolmoAct2 SO-101 Zero-Shot Evaluation Traces
This repository contains evaluation traces from running allenai/MolmoAct2-SO100_101 zero-shot on an SO-101 robot arm using the official LeRobot MolmoAct2 integration plus a remote async inference setup.
This is an evaluation artifact, not a training dataset or model checkpoint. The model under test is AllenAI's released MolmoAct2 SO-100/SO-101 checkpoint.
Summary
MolmoAct2 remote inference was successfully brought up on… See the full description on the dataset page: https://huggingface.co/datasets/abdul004/molmoact2-so101-zero-shot-eval.
