the.ai

Planning / Regimes

verified

Sim-to-Real

Simulators are fast, safe and free, and they are all wrong in ways that matter. A policy trained in one usually fails on the physical system it was meant for, because it learned details of the simulation rather than of the task. The counter-intuitive fix is not a better simulator but a deliberately worse one — many bad simulators instead of one good one.

Viz primitive · update-spectrumrandomisation = 0.2 · bars = 16

16 values. The left group decays steeply; the right group is 36% of the way to flat, and reads flatter than the left.

How much the policy's behaviour depends on each simulator parameter, before randomisation and after. Drag the randomisation up to watch the dependence even out — a policy that leans on nothing in particular is one that survives a system it was not trained on.

0.2

Reviewed by opendroid · 2026-08-18

  • arXiv:1703.06907 — Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World