Policy gradient
Policy-gradient methods adjust a parameterized action distribution to increase expected return. REINFORCE uses sampled returns to weight log-probability gradients.
Definition
Policy-gradient methods adjust a parameterized action distribution to increase expected return. REINFORCE uses sampled returns to weight log-probability gradients.
Intuition
Increase the probability of actions that performed better than expected, and reduce it for worse outcomes.
Worked example
If choosing a worked example earns more reward than the baseline, its sampled log-probability receives a positive learning signal.
The math
for an action-independent baseline.
In practice
Use fresh on-policy samples or a justified off-policy correction; replaying old log probabilities as if current can invalidate the update.
Go deeper
Sources
More in Reinforcement learning
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.