the.ai

Reinforcement / Methods

verified

Reward Shaping

Add extra rewards along the way so the agent gets feedback before the end. A robot rewarded only for reaching the goal learns nothing until it stumbles there by accident; rewarded a little for getting closer, it learns immediately. The obvious danger is equally immediate — reward getting closer and it may learn to hover just short of the goal forever.

Viz primitive · budget-splitnon-telescoping = 6

non-telescoping holds 13% of the budget; rest holds the remaining 87%.

Shaping reward that survives summing along a trajectory, against the part that cancels, in equal units. Drag the non-telescoping share up to watch the shaping start reordering trajectories — which is the same thing as changing which policy is optimal.

6

Reviewed by opendroid · 2026-08-18