Reinforcement learning

Contextual bandit

A contextual bandit chooses an action using the current context and receives one-step reward. It does not model how the action changes future states.

Ask the Reinforcement learning assistant1 min read · Updated September 9, 2026

Definition

A contextual bandit chooses an action using the current context and receives one-step reward. It does not model how the action changes future states.

Intuition

Choosing an answer format is often closer to a bandit than to a long-horizon control problem, unless future conversation effects are explicitly rewarded.

Worked example

A tutor selects a short explanation or a worked example for a question, then receives an evaluation score for that response.

The math

Choose a∼π(⋅∣x)a\sim\pi(\cdot|x) and optimize E[r(x,a)]\mathbb{E}[r(x,a)]. There is no bootstrapped next-state target.

In practice

Use a bandit for small routing decisions and compare it against rules and supervised learning before adding complexity.

Sources

More in Reinforcement learning

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.