RESEARCH PAPER

RoboGate: Adaptive Failure Discovery for Safe Robot Policy Deployment via Two-Stage Boundary-Focused Sampling

AgentAI Co., Ltd.|March 2026|DOI: 10.5281/zenodo.19166967

CORRECTION 2026-07-18 · LEARNED-POLICY COMPARISON RETRACTED

A subsequent audit found that the paper's VLA comparison used non-equivalent harness paths and an insufficient success predicate. Every associated run is quarantined and must not be cited as evidence of policy capability, safety or cross-simulator transfer. The scripted-controller failure-boundary study and the learned-policy benchmark are treated separately.

Read the full validity notice →

Publication Status

arXiv✅ Publishedcs.RO (2603.22126) v4 — 2026-04-20
Zenodo DOI: 10.5281/zenodo.19166967
SSRN Submitted — Under Review (Abstract ID: 6455499)

Abstract

This page describes the historical March 2026 paper on scripted-controller failure discovery. Correction dated 2026-07-18: the paper's learned-policy comparison and cross-simulator interpretation are retracted. A later audit found non-equivalent harness paths and an invalid learned-policy success predicate; the associated model runs are quarantined and must not be cited as capability results. The scripted-controller research dataset remains a separate historical experiment.

robot safetydeployment validationfailure analysisadaptive samplingsim-to-realVLA evaluation

Key Findings

Main results from 50,000+ Isaac Sim experiments across 4 robots

0.780Risk Model AUC

Combined two-stage model outperforms Stage 1 alone (0.754) by +3.4%

μ*(m)Boundary Equation

Closed-form: μ*(m) = (1.469 + 0.419m) / (3.691 - 1.400m) separates PASS/FAIL regions

z = -10.00friction × mass

Strongest interaction effect — failure cascades from timeout → collision → grasp miss

4Universal Danger Zones

mass > 0.93kg, friction < 0.492, friction×mass interaction, mass > 1.8kg → both robots fail

0% dropSuction Gripper

UR5e suction gripper eliminates drop failures entirely vs. Franka parallel-jaw

RETRACTEDLearned-policy comparison

The historical VLA-versus-scripted comparison used non-equivalent harness paths and is not valid benchmark evidence.

VLA Evaluation Correction

QUARANTINED · NO VALID SCORE

Raw run files remain available for transparency and forensic reproduction. The historical learned-policy results did not pass harness qualification, adapter preflight, complete provenance and transition-based scoring, so the paper's comparison table and conclusions are withdrawn. No replacement numeric comparison will be shown until qualified Harness v2 results exist.

Citation

If you use RoboGate in your research, please cite:

@misc{kim2026robogate,
  title         = {ROBOGATE: Adaptive Failure Discovery for Safe Robot
                   Policy Deployment via Two-Stage Boundary-Focused
                   Sampling},
  author        = {{AgentAI Co., Ltd.}},
  year          = {2026},
  eprint        = {2603.22126},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  doi           = {10.5281/zenodo.19166967},
  url           = {https://robogate.io/paper}
}