CORRECTION 2026-07-18 · LEARNED-POLICY COMPARISON RETRACTED
A subsequent audit found that the paper's VLA comparison used non-equivalent harness paths and an insufficient success predicate. Every associated run is quarantined and must not be cited as evidence of policy capability, safety or cross-simulator transfer. The scripted-controller failure-boundary study and the learned-policy benchmark are treated separately.
Read the full validity notice →| arXiv | ✅ Published — cs.RO (2603.22126) v4 — 2026-04-20 |
| Zenodo | ✅ DOI: 10.5281/zenodo.19166967 |
| SSRN | ✅ Submitted — Under Review (Abstract ID: 6455499) |
This page describes the historical March 2026 paper on scripted-controller failure discovery. Correction dated 2026-07-18: the paper's learned-policy comparison and cross-simulator interpretation are retracted. A later audit found non-equivalent harness paths and an invalid learned-policy success predicate; the associated model runs are quarantined and must not be cited as capability results. The scripted-controller research dataset remains a separate historical experiment.
Main results from 50,000+ Isaac Sim experiments across 4 robots
Combined two-stage model outperforms Stage 1 alone (0.754) by +3.4%
Closed-form: μ*(m) = (1.469 + 0.419m) / (3.691 - 1.400m) separates PASS/FAIL regions
Strongest interaction effect — failure cascades from timeout → collision → grasp miss
mass > 0.93kg, friction < 0.492, friction×mass interaction, mass > 1.8kg → both robots fail
UR5e suction gripper eliminates drop failures entirely vs. Franka parallel-jaw
The historical VLA-versus-scripted comparison used non-equivalent harness paths and is not valid benchmark evidence.
QUARANTINED · NO VALID SCORE
Raw run files remain available for transparency and forensic reproduction. The historical learned-policy results did not pass harness qualification, adapter preflight, complete provenance and transition-based scoring, so the paper's comparison table and conclusions are withdrawn. No replacement numeric comparison will be shown until qualified Harness v2 results exist.
If you use RoboGate in your research, please cite:
@misc{kim2026robogate,
title = {ROBOGATE: Adaptive Failure Discovery for Safe Robot
Policy Deployment via Two-Stage Boundary-Focused
Sampling},
author = {{AgentAI Co., Ltd.}},
year = {2026},
eprint = {2603.22126},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
doi = {10.5281/zenodo.19166967},
url = {https://robogate.io/paper}
}