Markov decision process
An MDP describes states, actions, transition probabilities, rewards, and a discount factor. A state is Markov when it contains the information needed to predict the next transition given the action.
Definition
An MDP describes states, actions, transition probabilities, rewards, and a discount factor. A state is Markov when it contains the information needed to predict the next transition given the action.
Intuition
A position alone is insufficient for a moving vehicle: velocity also affects what happens next. State design matters as much as the learning algorithm.
Worked example
In a grid world, the state can be the current cell, actions are four directions, and reaching the goal ends the episode.
The math
An MDP is commonly written . The transition model is .
In practice
Use the MDP description to test whether observation history or hidden variables are missing from your agent input.
Go deeper
Sources
More in Reinforcement learning
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.