Reinforcement learning

Experience replay

Experience replay stores transitions and samples them later for learning. It can reduce sample correlation and reuse expensive experience.

Ask the Reinforcement learning assistant1 min read · Updated September 9, 2026

Definition

Experience replay stores transitions and samples them later for learning. It can reduce sample correlation and reuse expensive experience.

Intuition

Reusing data saves environment interaction, but old data was collected under older behavior policies.

Worked example

A DQN buffer stores state, action, reward, next state, and terminal status, then samples a minibatch.

The math

Off-policy value learning can use replay under its assumptions; a naive on-policy REINFORCE update cannot simply reuse stale policy log probabilities.

In practice

Bound replay memory on a laptop and measure useful updates per second rather than maximizing buffer size.

Sources

More in Reinforcement learning

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.