Reinforcement learning

Value function

A state-value function estimates expected return from a state when following a particular policy. It depends on both the environment and the policy.

Ask the Reinforcement learning assistant1 min read · Updated September 9, 2026

Definition

A state-value function estimates expected return from a state when following a particular policy. It depends on both the environment and the policy.

Intuition

Value is the promise of future reward from where you are, averaged over what the policy will do.

Worked example

A cell near the goal can have low value if a hazardous transition makes failure likely. Distance alone does not determine value.

The math

Vπ(s)=Eπ[Gt∣St=s]V^\pi(s)=\mathbb{E}_{\pi}[G_t|S_t=s]. An action-value function additionally conditions on the first action.

In practice

Value estimates guide planning, actor-critic baselines, and debugging of unexpectedly preferred states.

Sources

More in Reinforcement learning

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.