CoolFace
Datasetpublic

mafangniu/Behavior-Skill

Behavior-Skill Behavior-Skill is a skill-centric dataset and evaluation benchmark built on BEHAVIOR-1K for Vision-Language-Action (VLA) policies in long-horizon mobile manipulation tasks. It establishes executable constituent skills as the fundamental unit for both policy learning and evaluation. Paper: arXiv  |  Code: GitHub Behavior-Skill contains 235,492 skill instances constructed from 10,000 demonstrations across 50 household tasks and 34 semantic skill… See the full description on the dataset page: https://huggingface.co/datasets/mafangniu/Behavior-Skill.

sourceHugging Facemitupdated 16d agoView on Hugging Face
1likes1.6kdownloads
Dataset Card

Behavior-Skill

Behavior-Skill is a skill-centric dataset and evaluation benchmark built on BEHAVIOR-1K for Vision-Language-Action (VLA) policies in long-horizon mobile manipulation tasks. It establishes executable constituent skills as the fundamental unit for both policy learning and evaluation.

Paper: arXiv  |  Code: GitHub

Behavior-Skill contains 235,492 skill instances constructed from 10,000 demonstrations across 50 household tasks and 34 semantic skill categories. Each annotated instance contains a skill instruction and its aligned frame interval in the corresponding BEHAVIOR-1K demonstration. For 500 evaluation demonstrations (10 per task), the release additionally provides per-skill BDDL success conditions and restorable OmniGibson states for independent closed-loop evaluation.

The original observations and action trajectories are available from the upstream BEHAVIOR-1K demonstrations and are referenced by frame intervals; they are not duplicated in this repository.

Dataset at a Glance

StatisticValue
Source benchmarkBEHAVIOR-1K
Household tasks50
Demonstrations10,000
Skill instances235,492
Semantic skill categories34
Average skill duration16.6 s
Evaluation demonstrations500 (10 per task)

Repository Contents

PathContentsPrimary use
skill_annotations/Episode-level skill instructions and aligned frame intervals for the complete annotation set.Skill policy training and data analysis
skill_eval_configs/Episode-level YAML files containing BDDL context and a success goal for every evaluated skill.Independent skill evaluation
skill_init_states/Restorable intermediate OmniGibson states and per-episode manifests for the evaluation set.Initializing skill-level rollouts under valid preconditions
eval_episodes.jsonMapping from the 50 tasks to their prompts, task IDs, and 10 evaluation episodes.Resolving the evaluation subset and file paths

The annotations cover the complete demonstration set. Evaluation configurations and initial states cover only the 500 evaluation demonstrations, not all 235,492 annotated skill instances.

Directory Structure

text
Behavior-Skill/
├── README.md
├── LICENSE
├── CITATION.cff
├── eval_episodes.json
├── skill_annotations/
│   ├── task-0000/
│   │   ├── episode_00000010_step.json
│   │   └── ...
│   └── task-0049/
├── skill_eval_configs/
│   ├── turning_on_radio/
│   │   ├── episode_00000010.yaml
│   │   └── ...
│   └── ...
└── skill_init_states/
    ├── task-0000/
    │   ├── episode_00000010/
    │   │   ├── skill_00_scene.json
    │   │   ├── skill_00_state.npz
    │   │   ├── ...
    │   │   └── state_manifest.json
    │   └── ...
    └── task-0049/

Data Formats

Skill annotations

Each skill_annotations/task-XXXX/episode_XXXXXXXX_step.json file describes the ordered constituent skills in one demonstration.

FieldDescription
taskTask name
task_descriptionTask-level instruction
cot_skill_listOrdered skill instructions
cot_skill_frame_durationOrdered start and end frames aligned one-to-one with cot_skill_list

Evaluation configurations

Each skill_eval_configs/<task_name>/episode_XXXXXXXX.yaml file contains task-level BDDL context and an ordered skills list.

FieldDescription
task_nameUnderscore-separated task name
activity_definition_idBEHAVIOR-1K activity definition ID
objects_bddlTyped BDDL object declarations
init_bddlRelevant initial BDDL facts
extra_object_name_mapOptional BDDL-to-scene mappings for objects absent from the original task BDDL
skillsSkill-level instructions, categories, types, and BDDL goals

Each skill entry contains id, instruction, skill_description(the semantic skill category), type, and bddl_goal.

Skill initial states

Each evaluation episode has a directory under skill_init_states/task-XXXX/episode_XXXXXXXX/.

FileDescription
skill_XX_scene.jsonRestorable OmniGibson scene snapshot for the state immediately before skill XX.
skill_XX_state.npzAlignment metadata containing skill, start_frame, and end_frame.
state_manifest.jsonMaps every skill index and frame interval to its matching scene and NPZ files and records the source HDF5 path.

State restoration is version-sensitive. The snapshots were generated with OmniGibson 3.7.2, BDDL 3.7.0, and BEHAVIOR-1K assets 3.7.2rc1; follow the environment specified by the accompanying code release.

Cross-Component Mapping

Task ID, episode ID, and local skill index connect the three components. For example, task-0000, episode 00000010, skill 02 maps to:

ComponentLocation
Annotationskill_annotations/task-0000/episode_00000010_step.jsoncot_skill_list[2]
BDDL goalskill_eval_configs/turning_on_radio/episode_00000010.yamlskills[2]
State manifestskill_init_states/task-0000/episode_00000010/state_manifest.jsonskills[2]
State filesskill_02_scene.json and skill_02_state.npz in the corresponding episode directory

Download

bash
pip install -U huggingface_hub

hf download mafangniu/Behavior-Skill \
  --repo-type dataset \
  --local-dir ./Behavior-Skill

To download a single component:

bash
hf download mafangniu/Behavior-Skill \
  --repo-type dataset \
  --include "skill_eval_configs/**" \
  --local-dir ./Behavior-Skill

Use a tagged release or fixed Hub revision for reproducible experiments.

License

The original Behavior-Skill annotations, evaluation configurations, state snapshots, metadata, and documentation are released under the MIT License. Third-party datasets, simulators, and assets remain subject to their respective licenses and terms. See LICENSE for details.

Citation

If you find Behavior-Skill useful in your research, please cite our paper:

bibtex
@article{ma2026behaviorskill,
  title   = {{Behavior-Skill}: A Fine-Grained Benchmark for Evaluating Vision-Language-Action Policies in Long-Horizon Tasks},
  author  = {Ma, Chunyun and Luo, Lun and Luo, Xingjian and Feng, Xiexing and Zhang, Hang and Liu, Wei and Qiao, Feng and Wang, Yaonan and Lu, Huimin and Chen, Xieyuanli},
  journal = {arXiv preprint arXiv:2608.30536},
  year    = {2026},
  url     = {https://arxiv.org/abs/2608.30536}
}

Please also cite the original BEHAVIOR-1K benchmark:

bibtex
@inproceedings{li2023behavior,
  title={Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation},
  author={Li, Chengshu and Zhang, Ruohan and Wong, Josiah and Gokmen, Cem and Srivastava, Sanjana and Mart{\'\i}n-Mart{\'\i}n, Roberto and Wang, Chen and Levine, Gabrael and Lingelbach, Michael and Sun, Jiankai and others},
  booktitle={Conference on Robot Learning},
  pages={80--93},
  year={2023},
  organization={PMLR}
}

Contact

For questions or reproducible data issues, open an issue in the Behavior-Skill code repository or contact Chunyun Ma at mafangniu@gmail.com.