Reinforcement learning

Markov decision process

An MDP describes states, actions, transition probabilities, rewards, and a discount factor. A state is Markov when it contains the information needed to predict the next transition given the action.

Ask the Reinforcement learning assistant 1 min read · Updated September 9, 2026

Definition

An MDP describes states, actions, transition probabilities, rewards, and a discount factor. A state is Markov when it contains the information needed to predict the next transition given the action.

Intuition

A position alone is insufficient for a moving vehicle: velocity also affects what happens next. State design matters as much as the learning algorithm.

Worked example

In a grid world, the state can be the current cell, actions are four directions, and reaching the goal ends the episode.

The math

An MDP is commonly written (S,A,P,R,γ)(S,A,P,R,\gamma). The transition model is P(s′∣s,a)P(s^{\prime}|s,a).

In practice

Use the MDP description to test whether observation history or hidden variables are missing from your agent input.

Go deeper

Sources

More in Reinforcement learning

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.