16 topics
Reinforcement learning
Rewards, value functions, bandits, Q-learning and policy gradients.
A
B
C
E
- Efficient RL on a laptopEfficient laptop training starts with a small decision problem, compact features, vectorized batches, and a fixed wall-clock budget.
- Experience replayExperience replay stores transitions and samples them later for learning. It can reduce sample correlation and reuse expensive experience.
- Exploration and exploitationExploration tries actions to learn about their outcomes; exploitation chooses actions that currently appear best.
M
P
Q
R
- Reinforcement learningReinforcement learning studies agents that improve decisions through interaction and reward. The objective is expected return over time…
- Reward and returnReward is feedback for one transition; return aggregates rewards across future steps. They are different learning targets.
- Reward shapingReward shaping adds learning guidance. Poorly designed rewards can teach an agent to optimize a proxy while failing the intended task.
- RL evaluationRL evaluation measures a fixed policy on tasks, seeds, or environments that were not used to choose its parameters. Training reward alone…