the.ai

Reinforcement / Methods

verified

Temporal Difference Learning

Wait until the end of the game to learn what your moves were worth and you learn slowly and rarely. Temporal difference learning updates on every step instead, by comparing what you predicted with what you predicted one step later — bootstrapping off your own estimate rather than waiting for the truth.

Viz primitive · budget-splitbootstrapped-steps = 1

bootstrapped-steps holds 5% of the budget; rest holds the remaining 95%.

Steps whose value comes from the agent's own estimate against those backed by observed reward. Drag the bootstrapping up to trade the variance of waiting for the bias of guessing.

1

Reviewed by opendroid · 2026-08-17

Origin · not linkable

  • Sutton 1988 — Learning to Predict by the Methods of Temporal Differences · Machine Learning 3(1) · doi:10.1007/BF00115009