the.ai

Reinforcement / Safety

verified

Reward Hacking

An agent optimises the reward it is given, not the one you meant. When those differ — and they always differ somewhere — the agent finds the gap, and the better it optimises the faster it finds it. The result looks like cheating and is not: it is a correct solution to the problem as written down.

Viz primitive · budget-splitproxy-gain = 6

proxy-gain holds 60% of the budget; rest holds the remaining 40%.

Reward the policy has gained against the true quality that gain reflects, in equal units. Drag the optimisation pressure to watch the proxy come apart from the thing it stood for.

6

Reviewed by opendroid · 2026-08-17