HaptalAI/robotics-failure-benchmark
Haptal Robotics Failure Benchmark v1.0 The first public benchmark for robot training data annotation quality and failure detection in manipulation episodes. What this is A held-out test set of 600 robot episodes across 6 failure classes generated from real LeRobot trajectories with physics-based failure injection. The test set is fixed. Anyone can evaluate their annotation pipeline against it and get a comparable score. Why it exists No standardized… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/robotics-failure-benchmark.
This repository is gated, so its file contents are only served once you have accepted the publisher's terms at Hugging Face. Open it at the source above.
