CoolFace
Datasetpublic

BAAI/Discoverse-L

This dataset is the result of joint research conducted by Hao Tang's team at the School of Computer Science, Peking University, and the Beijing Academy of Artificial Intelligence (BAAI). Paper: EvoVLA: Self-Evolving Vision-Language-Action Model Authors: Zeting Liu*, Zida Yang*, Zeyu Zhang*†, Hao Tang‡Institution: Peking University Overview Discoverse-L is a long-horizon manipulation benchmark built on the DISCOVERSE simulator with AIRBOT-Play robot platform. It… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Discoverse-L.

sourceHugging Faceupdated 21d agoView on Hugging Face
2likes694downloads
Dataset Card

This dataset is the result of joint research conducted by Hao Tang's team at the School of Computer Science, Peking University, and the Beijing Academy of Artificial Intelligence (BAAI).

Paper: EvoVLA: Self-Evolving Vision-Language-Action Model Authors: Zeting Liu, Zida Yang, Zeyu Zhang†, Hao Tang‡ Institution*: Peking University

Overview

Discoverse-L is a long-horizon manipulation benchmark built on the DISCOVERSE simulator with AIRBOT-Play robot platform. It provides:

  • —3 multi-stage manipulation tasks with varying difficulty:
  • —Block Bridge (74 stages): Place two bars to form a bridge structure, then fill with multiple blocks
  • —Stack (18 stages): Stack three colored blocks in sequence
  • —Jujube-Cup (19 stages): Place a jujube fruit into a cup and move the cup onto a plate
  • —50 scripted demonstration trajectories per task (150 total)
  • —Task-aligned normalization statistics for cross-task training
  • —Stage dictionaries with Gemini-generated triplets (positive, negative, hard-negative text descriptions)

Dataset Structure

Discoverse-L/
├── demonstrations/
│   ├── block_bridge_place/
│   │   ├── 000/
│   │   │   ├── obs_action.json    # Joint states & actions
│   │   │   ├── cam_0.mp4           # Main camera view
│   │   │   └── cam_1.mp4           # Wrist camera view
│   │   ├── 001/
│   │   └── ... (50 episodes)
│   ├── place_jujube_coffeecup/
│   │   └── ... (50 episodes)
│   └── stack_block/
│       └── ... (50 episodes)
├── metadata/
│   └── task_aligned_normalization.json  # q01/q99/mean/std for each task
└── stage_dictionaries/
    ├── block_bridge_place_stages.json
    ├── place_jujube_coffeecup_stages.json
    └── stack_block_stages.json

Data Format

Demonstration Trajectories

Each episode directory contains:

  • —obs_action.json: Time-aligned observations and actions
json
  {
    "time": [t1, t2, ...],
    "obs": {
      "jq": [[q0, q1, q2, q3, q4, q5, q6], ...]  // Joint positions
    },
    "act": [[a0, a1, a2, a3, a4, a5, a6], ...]   // Joint actions
  }
  • —cam_0.mp4: 448×448 main camera view (20 FPS)
  • —cam_1.mp4: 448×448 wrist camera view (20 FPS)

Task-Aligned Normalization

Computed from the 50 demonstrations per task:

json
{
  "task_name": {
    "action": {
      "mean": [7-dim],
      "std": [7-dim],
      "min": [7-dim],
      "max": [7-dim],
      "q01": [7-dim],  // 1st percentile
      "q99": [7-dim],  // 99th percentile
      "mask": [bool×7] // True for continuous joints, False for gripper
    }
  }
}

Stage Dictionaries

Gemini-2.5-Pro generated text triplets for each stage:

json
[
  {
    "id": 0,
    "positive": "The robotic gripper is approaching the target object",
    "negative": "The gripper is moving away from all objects",
    "hard_negative": "The gripper is grasping a distractor object"
  },
  ...
]
python
from datasets import load_dataset

load_dataset("BAAI/Discoverse-L")