CoolFace
Datasetpublic

d3d3shan/pi05-libero-plus-language-instructions-failures

pi0.5 LIBERO-plus Language Instructions Failures Failure summary artifacts for evaluating TensorAuto/tPi0.5-libero on the LIBERO-plus Language Instructions perturbation subset through OpenTau. Evaluation Setup Benchmark: LIBERO-plus Suite: libero_10 Perturbation category: Language Instructions Policy: TensorAuto/tPi0.5-libero Tasks: 383 Episodes per task: 5 Metric episodes: 1915 Seed schedule: episode seeds 1000 to 1004 Episode length: 520 steps Source machine:… See the full description on the dataset page: https://huggingface.co/datasets/d3d3shan/pi05-libero-plus-language-instructions-failures.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes72downloads
Dataset Card

pi0.5 LIBERO-plus Language Instructions Failures

Failure summary artifacts for evaluating TensorAuto/tPi0.5-libero on the LIBERO-plus Language Instructions perturbation subset through OpenTau.

Evaluation Setup

  • —Benchmark: LIBERO-plus
  • —Suite: libero_10
  • —Perturbation category: Language Instructions
  • —Policy: TensorAuto/tPi0.5-libero
  • —Tasks: 383
  • —Episodes per task: 5
  • —Metric episodes: 1915
  • —Seed schedule: episode seeds 1000 to 1004
  • —Episode length: 520 steps
  • —Source machine: wzxuan under /data/wangjinqiu/tta-memory/libero_plus_eval

Summary

  • —Successes: 1841 / 1915
  • —Success rate: 96.14%
  • —Failures: 74
  • —Failure grid videos: 5
  • —Failures missing videos in the grid report: 0

Files

  • —videos/: 16-panel failure grid videos.
  • —metadata/summary.json: benchmark counts, success rate, and difficulty-level breakdown.
  • —metadata/metric_records.json: per-episode metric records used for the success rate.
  • —metadata/eval_info_paths.json: eval files included in the payload.
  • —metadata/source_configs/: configs/logs copied from the wzxuan run payload when available.

The grid videos are for qualitative failure-mode inspection. The official success rate above is computed only from metric eval outputs and excludes failed-task video reruns.