Contextual bandit
A contextual bandit chooses an action using the current context and receives one-step reward. It does not model how the action changes future states.
Definition
A contextual bandit chooses an action using the current context and receives one-step reward. It does not model how the action changes future states.
Intuition
Choosing an answer format is often closer to a bandit than to a long-horizon control problem, unless future conversation effects are explicitly rewarded.
Worked example
A tutor selects a short explanation or a worked example for a question, then receives an evaluation score for that response.
The math
Choose and optimize . There is no bootstrapped next-state target.
In practice
Use a bandit for small routing decisions and compare it against rules and supervised learning before adding complexity.
Sources
More in Reinforcement learning
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.