CoolFace
Datasetpublic

i-am-shaurya05/robotrace-vla-robustness-traces

RoboTrace Evidence Bundle This dataset repository contains the public evidence bundle for RoboTrace, a low-cost deployment-stress evaluation scaffold for robot-learning and VLA-style inference pipelines. The current release evaluates lerobot/pusht and includes reports, metrics, plots, summaries, and release manifests from a complete staged run. What this bundle is for Use this repository to inspect evidence from RoboTrace: action-trace stability metrics visual… See the full description on the dataset page: https://huggingface.co/datasets/i-am-shaurya05/robotrace-vla-robustness-traces.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes70downloads
Dataset Card

RoboTrace Evidence Bundle

This dataset repository contains the public evidence bundle for RoboTrace, a low-cost deployment-stress evaluation scaffold for robot-learning and VLA-style inference pipelines.

The current release evaluates lerobot/pusht and includes reports, metrics, plots, summaries, and release manifests from a complete staged run.

What this bundle is for

Use this repository to inspect evidence from RoboTrace:

  • —action-trace stability metrics
  • —visual perturbation severity results
  • —symbolic instruction perturbation artifacts
  • —async deployment-stress simulations
  • —lightweight baseline summaries
  • —limitations and reproducibility notes

Why it matters

RoboTrace is relevant to edge AI and inference deployment because robot/VLA systems can fail from runtime issues even when the model itself is not being evaluated directly.

This includes stale observations, delayed frames, reduced control frequency, action chunk reuse, and action-horizon mismatch.

Run summary

ItemValue
Datasetlerobot/pusht
Splittrain
Usable episodes4
Step metric rows512
Visual frames sampled12
Visual perturbation rows144
Async stress scenarios26
Async simulation rows104
Baselines evaluated8
Baseline eval rows32
Real VLA policy loadedFalse

Main finding

The strongest deployment-stress failure mode in this run was action-horizon mismatch.

Highest-risk scenario:

  • —Mode: action_horizon_mismatch
  • —Params: {"actual_horizon": 8, "expected_horizon": 4}
  • —Risk score proxy: 2.5562635495893855
  • —Mean L2 drift: 138.2900505065918

Repository contents

  • —artifacts/reports/: final report, public summary, limitations, reproducibility notes
  • —artifacts/summaries/: CSV/JSON metrics and run summaries
  • —artifacts/figures/: generated plots
  • —manifest/: release manifest and file index
  • —release_notes.md: release notes
  • —publish_instructions.md: publishing notes

Key files

  • —artifacts/reports/robotrace_report.md
  • —artifacts/reports/robotrace_public_summary.md
  • —artifacts/reports/limitations.md
  • —artifacts/reports/reproducibility.md
  • —artifacts/summaries/robotrace_final_metrics.csv
  • —artifacts/summaries/robotrace_final_summary.json
  • —artifacts/summaries/robotrace_evidence_index.csv

Scope and limitations

This release does not claim real robot hardware validation, real VLA policy robustness, natural-language instruction robustness for lerobot/pusht, task success-rate degradation, measured serving latency, or SOTA benchmark performance.

RoboTrace is an offline trace-level and simulation-level evaluation scaffold for producing reproducible engineering evidence under low-cost compute constraints.

Related links

  • —GitHub: https://github.com/githubshaurya/RoboTrace
  • —Space dashboard: https://huggingface.co/spaces/i-am-shaurya05/robotrace-vla-dashboard