CoolFace
Datasetpublic

liveplex/robogate-failure-dictionary

RoboGate Failure Dictionary 50,000+ Physics-Validated Pick & Place Failure Patterns across 4 Robots (Franka Panda, UR5e, UR3e, UR10e) A structured database of robot AI failure patterns collected from NVIDIA Isaac Sim physical simulations using Two-Stage Adaptive Sampling. Each experiment records the exact conditions under which a robot succeeded or failed at Pick & Place tasks. Quick Stats Franka Uniform Franka Boundary UR5e UR3e UR10e Combined… See the full description on the dataset page: https://huggingface.co/datasets/liveplex/robogate-failure-dictionary.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes156downloads
README.md142 linesDownload Raw Back to root
1---2license: mit3task_categories:4  - robotics5tags:6  - failure-analysis7  - pick-and-place8  - isaac-sim9  - franka-panda10  - ur5e11  - ur3e12  - ur10e13  - domain-randomization14  - latin-hypercube-sampling15  - adaptive-sampling16  - physical-ai17size_categories:18  - 10K<n<100K19language:20  - en21  - ko22---23 24# RoboGate Failure Dictionary25 26> **50,000+ Physics-Validated Pick & Place Failure Patterns across 4 Robots (Franka Panda, UR5e, UR3e, UR10e)**27 28A structured database of robot AI failure patterns collected from **NVIDIA Isaac Sim** physical simulations using **Two-Stage Adaptive Sampling**. Each experiment records the exact conditions under which a robot succeeded or failed at Pick & Place tasks.29 30## Quick Stats31 32| | Franka Uniform | Franka Boundary | UR5e | UR3e | UR10e | Combined |33|---|---|---|---|---|---|---|34| **Experiments** | 10,000 | 10,000 | 10,000 | 10,000 | 10,000 | **50,000+** |35| **Success Rate** | 33.3% | 63.8% | 74.3% | 10.0% | 0.0% | — |36| **Franka Combined** | — | 48.6% | — | — | — | — |37| **Risk Model AUC** | 0.65 | **0.777** | — | — | — | 0.777 |38| **Sampling** | Uniform LHS | Boundary LHS | Uniform LHS | Uniform LHS | Uniform LHS | Two-Stage |39 40## Key Findings41 42- **friction × mass interaction z = -10.00** — strongest predictor of failure43- **Friction threshold: 0.492 ± 0.031** — below this, failure cascades44- **Mass > 0.93 kg** → Both robots fail at **< 40%** SR (universal danger zone)45- **Boundary equation:** μ*(m) = (1.469 + 0.419m) / (3.691 - 1.400m)46- **AUC improved 0.65 → 0.777** (+19.5%) with boundary-focused sampling47- **Failure mode transition:** friction↓ → timeout → collision → grasp_miss48 49## Two-Stage Adaptive Sampling50 51**Stage 1 — Uniform Exploration (40,000)**52- Franka Panda 10K + UR5e 10K53- Latin Hypercube Sampling for uniform parameter space coverage54- Identified boundary regions and initial risk model (AUC 0.65)55 56**Stage 2 — Boundary-Focused (10,000)**57- Franka Panda only, targeting boundary/transition regions58- Concentrated sampling near friction threshold 0.49259- Revealed failure mode transitions invisible to uniform sampling60- Boosted Risk Model AUC to 0.777 (+19.5%)61 62## Universal Danger Zones (mass > 0.93 kg)63 64| Mass Range | Franka SR | UR5e SR |65|---|---|---|66| 0.93 – 1.23 kg | 21.4% | 30.9% |67| 1.23 – 1.52 kg | 14.9% | 25.3% |68| 1.52 – 1.82 kg | 12.5% | 28.9% |69| 1.82 – 2.11 kg | 6.6% | 28.1% |70 71## Usage72 73```python74from datasets import load_dataset75 76ds = load_dataset("liveplex/robogate-failure-dictionary")77print(ds["train"][0])78 79# Filter danger zones80danger = ds["train"].filter(lambda x: x["zone"] == "danger")81print(f"Danger zones: {len(danger)}")82```83 84## Parameter Space85 86| Parameter | Range | Scale | Paper |87|-----------|-------|-------|-------|88| friction | 0.05 – 1.2 | log-uniform | SIMPLER 2024 |89| mass | 0.05 – 2.0 kg | log-uniform | SIMPLER 2024 |90| com_offset | 0.0 – 0.40 | uniform | Suction Grasp 2025 |91| size | 0.02 – 0.12 m | uniform | SIMPLER 2024 |92| ik_noise | 0.0 – 0.04 rad | uniform | ICRA Sim2Real 2025 |93| obstacles | 0 – 4 | integer | RoboFAC 2025 |94| shape | 5 types | categorical | Grasp Anything 2024 |95| placement | 14 types | categorical | ALEAS 2025 |96 97## Research Foundations98 99| Design Choice | Paper | Year |100|---|---|---|101| Two-Stage Adaptive Sampling | ALEAS | 2025 |102| friction × mass interaction | SIMPLER | CoRL 2024 |103| Failure taxonomy | RoboFAC | NeurIPS 2025 |104| Cross-robot validation | RoboMIND | RSS 2025 |105| UR-specific failures | Guardian | ICRA 2025 |106| Confidence intervals | SureSim | Badithela 2025 |107| GPU simulation | Isaac Lab | NVIDIA 2025 |108| Grasp evaluation | Isaac Sim Grasping SDG | NVIDIA 2025 |109 110## VLA Benchmark — 4-Model Leaderboard111 112Four VLA models evaluated on RoboGate's 68-scenario adversarial suite. **All scored 0% SR** — including NVIDIA's official GR00T N1.6.113 114| Model | Params | SR | Confidence | Failure Pattern |115|-------|--------|-----|-----------|-----------------|116| Scripted Controller | — | **100%** (68/68) | 76/100 | — |117| **GR00T N1.6 (NVIDIA)** | 3B | 0% (0/68) | 1/100 | grasp_miss + collision |118| OpenVLA (Stanford + TRI) | 7B | 0% (0/68) | 27/100 | grasp_miss dominant, 0 collision |119| Octo-Base (UC Berkeley) | 93M | 0% (0/68) | 1/100 | grasp_miss 79%, collision 21% |120| Octo-Small (UC Berkeley) | 27M | 0% (0/68) | 1/100 | grasp_miss 79.4%, collision 20.6% |121 122Model size is not the bottleneck — even NVIDIA's flagship 3B model cannot bridge the training-deployment distribution gap.123 124**Leaderboard:** [robogate.io/vla](https://robogate.io/vla) · **Paper:** [arXiv:2603.22126](https://arxiv.org/abs/2603.22126)125 126## Citation127 128```bibtex129@dataset{robogate_failure_dictionary_2026,130  title={RoboGate Failure Dictionary: 50K+ Physics-Validated Pick & Place Failure Patterns},131  author={RoboGate Team},132  year={2026},133  url={https://huggingface.co/datasets/liveplex/robogate-failure-dictionary},134  note={Franka Panda + UR3e + UR5e + UR10e, Two-Stage Adaptive Sampling, AUC 0.777}135}136```137 138## Links139 140- **GitHub**: [liveplex-cpu/robogate-failure-dictionary](https://github.com/liveplex-cpu/robogate-failure-dictionary)141- **RoboGate Platform**: [liveplex-cpu/robogate](https://github.com/liveplex-cpu/robogate)142