Reinforcement learning

Bellman equation

A Bellman equation relates value now to expected immediate reward plus discounted value after the next transition.

Ask the Reinforcement learning assistant1 min read · Updated September 9, 2026

Definition

A Bellman equation relates value now to expected immediate reward plus discounted value after the next transition.

Intuition

It breaks a long-horizon prediction into a local consistency condition.

Worked example

With certain reward 2, next-state value 5, and discount 0.9, a consistent current value is 2 + 0.9 × 5 = 6.5.

The math

Vπ(s)=∑aπ(a∣s)∑s′P(s′∣s,a)[R(s,a,s′)+γVπ(s′)]V^\pi(s)=\sum_a\pi(a|s)\sum_{s^{\prime}}P(s^{\prime}|s,a)[R(s,a,s^{\prime})+\gamma V^\pi(s^{\prime})].

In practice

Use Bellman residuals to diagnose value estimates, while remembering that a small training residual does not guarantee good unseen behavior.

Sources

More in Reinforcement learning

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.