RoboGate turns simulation and field telemetry into reproducible failure evidence for a declared policy, robot, cell and operating envelope. Historical research data remain available, while public benchmark scores are published only after strict qualification.
GitHub Dataset
50,000+ experiments · 4 robots · MIT License
HuggingFace Dataset
robogate-failure-dictionary · Robotics
50,000+ Isaac Sim experiments · Franka Panda + UR5e + UR3e + UR10e · NVIDIA RTX
ILLUSTRATIVE OUTPUT — NOT A LIVE RESULT
$ robogate test --policy /path/to/policy --config test.yaml
[1/5] Verifying artifacts... ✓
[2/5] Loading config... ✓
[3/5] Running 68 scenarios...
████████████████████████ 68/68
[4/5] Computing metrics... ✓
[5/5] Generating reports... ✓✅ PASS — Confidence: 92/100
Success Rate: 95.4% ▲ +2.2%
Collisions: 0
Cycle Time: 4.1s
Reliability Claims Need Evidence That Survives Audit
A simulator run is useful only when the scenario, adapter, model, controls, state transitions and provenance are all verifiable.
Our July 2026 audit found that the previous fair-harness cohort could credit exogenous object motion as policy success. We quarantined all 20 historical model runs and paused automatic evaluation and publication.
Historical D-041 cohort
20
Qualified public results
0
Harness v2 now fails closed: no result is ranked until the harness, every scored scenario, the model adapter and the complete 68-episode artifact pass qualification.
RoboGate is a Physical AI reliability data engine — not a one-score safety certificate.
Current status: automatic evaluation and publication remain paused pending Harness v2 GPU qualification and a validated model canary. Read the 2026-07-18 validity notice.
50,000+ scripted-controller simulation records · NVIDIA Isaac Sim · four robot configurations · separate from the quarantined learned-policy benchmark
Experiments
Risk Model AUC
Friction Threshold
friction × mass interaction
NVIDIA built the simulator. RoboGate built deployment validation on top of it.
You only test under normal conditions. Dark environments, tiny objects, surfaces with 0.1 friction — your robot encounters these for the first time in production.
Edge case failure rate 83%+Senior engineers run tests manually. Criteria vary by person. Results are rarely documented.
RoboGate: 68 scenarios, fully automatedNo alerts when success rate drops by 5%. Problems surface only after production stops.
Drift detection: under 5 minRun versioned scenarios through one state-machine scorer with negative and positive controls.
manifest · immutable harness · adapter preflight · strict validator
Unqualified runs are ERROR or QUARANTINED — never silently ranked
Map failures to the exact policy, embodiment, cell and operating envelope.
raw evidence · provenance · revision history · actionable conditions
A simulator result is evidence for its declared scope, not a safety certificate
Ingest field telemetry and detect policy- and recipe-aware drift after deployment.
explicit baseline · cooldown · deduplication · health checks
Slack · Telegram · webhook integrations
Franka Panda performing Pick & Place tasks in Isaac Sim
good_policy — Validation Gate PASS
bad_policy — Validation Gate FAIL (collisions detected)
robogate test recorded live on our RTX 5090 host — real Isaac Sim, real reports, customer artifacts destroyed on completion
policy_v2 · nominal acceptance 20/20 · PASS · confidence 90.5
policy_v3 (untrained) · 0/20 · 20 collisions · FAIL · rollback recommended
Historical CLI recording from 2026-06-12. It demonstrates the legacy command flow only; its scripted-controller scores are not comparable to learned-policy runs and are not current benchmark evidence.
Define the policy, robot, cell and operating envelope; RoboGate returns scoped failure evidence rather than a blanket safety claim.
50,000+ Isaac Sim experiments completed · MIT License · NVIDIA Isaac Sim 5.1