d3d3shan/pi05-libero-plus-language-instructions-failures
pi0.5 LIBERO-plus Language Instructions Failures Failure summary artifacts for evaluating TensorAuto/tPi0.5-libero on the LIBERO-plus Language Instructions perturbation subset through OpenTau. Evaluation Setup Benchmark: LIBERO-plus Suite: libero_10 Perturbation category: Language Instructions Policy: TensorAuto/tPi0.5-libero Tasks: 383 Episodes per task: 5 Metric episodes: 1915 Seed schedule: episode seeds 1000 to 1004 Episode length: 520 steps Source machine:… See the full description on the dataset page: https://huggingface.co/datasets/d3d3shan/pi05-libero-plus-language-instructions-failures.
pi0.5 LIBERO-plus Language Instructions Failures
Failure summary artifacts for evaluating TensorAuto/tPi0.5-libero on the LIBERO-plus Language Instructions perturbation subset through OpenTau.
Evaluation Setup
- Benchmark: LIBERO-plus
- Suite:
libero_10 - Perturbation category:
Language Instructions - Policy:
TensorAuto/tPi0.5-libero - Tasks: 383
- Episodes per task: 5
- Metric episodes: 1915
- Seed schedule: episode seeds
1000to1004 - Episode length: 520 steps
- Source machine:
wzxuanunder/data/wangjinqiu/tta-memory/libero_plus_eval
Summary
- Successes: 1841 / 1915
- Success rate: 96.14%
- Failures: 74
- Failure grid videos: 5
- Failures missing videos in the grid report: 0
Files
videos/: 16-panel failure grid videos.metadata/summary.json: benchmark counts, success rate, and difficulty-level breakdown.metadata/metric_records.json: per-episode metric records used for the success rate.metadata/eval_info_paths.json: eval files included in the payload.metadata/source_configs/: configs/logs copied from the wzxuan run payload when available.
The grid videos are for qualitative failure-mode inspection. The official success rate above is computed only from metric eval outputs and excludes failed-task video reruns.
