Q-learning
Q-learning is an off-policy temporal-difference method that updates an action-value estimate toward a reward plus the best estimated next-state value.
Definition
Q-learning is an off-policy temporal-difference method that updates an action-value estimate toward a reward plus the best estimated next-state value.
Intuition
It learns about a greedy target policy even when behavior explores other actions.
Worked example
If Q = 2, reward = 1, next best Q = 4, discount = 0.9, and learning rate = 0.5, the updated value is 3.3.
The math
. At a true terminal state, omit the bootstrap term.
In practice
A small tabular Q-learning task is a useful first baseline before neural networks and replay buffers.
Go deeper
Sources
More in Reinforcement learning
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.