Reinforcement learning

Policy gradient

Policy-gradient methods adjust a parameterized action distribution to increase expected return. REINFORCE uses sampled returns to weight log-probability gradients.

Ask the Reinforcement learning assistant 1 min read · Updated September 9, 2026

Definition

Policy-gradient methods adjust a parameterized action distribution to increase expected return. REINFORCE uses sampled returns to weight log-probability gradients.

Intuition

Increase the probability of actions that performed better than expected, and reduce it for worse outcomes.

Worked example

If choosing a worked example earns more reward than the baseline, its sampled log-probability receives a positive learning signal.

The math

∇J(θ)=E[∇log⁡πθ(a∣s)(G−b(s))]\nabla J(\theta)=\mathbb{E}[\nabla\log\pi_\theta(a|s)(G-b(s))] for an action-independent baseline.

In practice

Use fresh on-policy samples or a justified off-policy correction; replaying old log probabilities as if current can invalidate the update.

Go deeper

Sources

More in Reinforcement learning

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.