Bellman equation
A Bellman equation relates value now to expected immediate reward plus discounted value after the next transition.
Definition
A Bellman equation relates value now to expected immediate reward plus discounted value after the next transition.
Intuition
It breaks a long-horizon prediction into a local consistency condition.
Worked example
With certain reward 2, next-state value 5, and discount 0.9, a consistent current value is 2 + 0.9 × 5 = 6.5.
The math
.
In practice
Use Bellman residuals to diagnose value estimates, while remembering that a small training residual does not guarantee good unseen behavior.
Sources
More in Reinforcement learning
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.