the.ai

Reinforcement / Foundations

verified

Credit Assignment

Something good happened. Which of the things you did made it happen? The question sounds simple and is the central difficulty of learning from delayed feedback — a reward arriving at move two hundred says nothing about which of the two hundred mattered, and most of them did not.

Viz primitive · budget-splittrajectory-length = 8

trajectory-length holds 29% of the budget; rest holds the remaining 71%.

Decisions too far back for one end-of-episode reward to speak to, against the recent ones it plausibly does, in decisions. Drag the trajectory length up to watch the unattributable share take over — the reward is not rarer, it is less informative about any one thing you did.

8

Reviewed by opendroid · 2026-08-18

  • arXiv:1806.07857 — RUDDER: Return Decomposition for Delayed Rewards
  • arXiv:1506.02438 — High-Dimensional Continuous Control Using Generalized Advantage Estimation