the.ai

Reinforcement / Methods

verified

Advantage Estimation

Was this action better than what I would normally have done here? That is the advantage — the value of an action minus the value of the state it was taken in — and it is a far more useful learning signal than the raw return, because it strips out how good the situation was and leaves only what the choice contributed.

Viz primitive · budget-splitvariance-share = 8

variance-share holds 17% of the budget; rest holds the remaining 83%.

Error from variance in the estimate, against error from the bias traded for it, in equal units. Drag the variance share up toward the Monte Carlo end to watch it dominate — neither endpoint minimises the total, which is why λ sits in between.

8

Reviewed by opendroid · 2026-08-18