Reinforcement learning

Q-learning

Q-learning is an off-policy temporal-difference method that updates an action-value estimate toward a reward plus the best estimated next-state value.

Ask the Reinforcement learning assistant1 min read · Updated September 9, 2026

Definition

Q-learning is an off-policy temporal-difference method that updates an action-value estimate toward a reward plus the best estimated next-state value.

Intuition

It learns about a greedy target policy even when behavior explores other actions.

Worked example

If Q = 2, reward = 1, next best Q = 4, discount = 0.9, and learning rate = 0.5, the updated value is 3.3.

The math

Q(s,a)←Q(s,a)+α[r+γmax⁡a′Q(s′,a′)−Q(s,a)]Q(s,a)\leftarrow Q(s,a)+\alpha[r+\gamma\max_{a^{\prime}}Q(s^{\prime},a^{\prime})-Q(s,a)]. At a true terminal state, omit the bootstrap term.

In practice

A small tabular Q-learning task is a useful first baseline before neural networks and replay buffers.

Go deeper

Sources

More in Reinforcement learning

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.