Skip to the content.

Week 05 — Reinforcement Learning II

← Reinforcement Learning I · Week 5 of 11 · Next: Generative Models →

Complete lecture deck

Complete deck: 36 slides rendered locally. Select any slide to open the full-resolution image.
Lecture 5: Reinforcement Learning II - slide 1 of 36
Slide 1 of 36
Lecture 5: Reinforcement Learning II - slide 2 of 36
Slide 2 of 36
Lecture 5: Reinforcement Learning II - slide 3 of 36
Slide 3 of 36
Lecture 5: Reinforcement Learning II - slide 4 of 36
Slide 4 of 36
Lecture 5: Reinforcement Learning II - slide 5 of 36
Slide 5 of 36
Lecture 5: Reinforcement Learning II - slide 6 of 36
Slide 6 of 36
Lecture 5: Reinforcement Learning II - slide 7 of 36
Slide 7 of 36
Lecture 5: Reinforcement Learning II - slide 8 of 36
Slide 8 of 36
Lecture 5: Reinforcement Learning II - slide 9 of 36
Slide 9 of 36
Lecture 5: Reinforcement Learning II - slide 10 of 36
Slide 10 of 36
Lecture 5: Reinforcement Learning II - slide 11 of 36
Slide 11 of 36
Lecture 5: Reinforcement Learning II - slide 12 of 36
Slide 12 of 36
Lecture 5: Reinforcement Learning II - slide 13 of 36
Slide 13 of 36
Lecture 5: Reinforcement Learning II - slide 14 of 36
Slide 14 of 36
Lecture 5: Reinforcement Learning II - slide 15 of 36
Slide 15 of 36
Lecture 5: Reinforcement Learning II - slide 16 of 36
Slide 16 of 36
Lecture 5: Reinforcement Learning II - slide 17 of 36
Slide 17 of 36
Lecture 5: Reinforcement Learning II - slide 18 of 36
Slide 18 of 36
Lecture 5: Reinforcement Learning II - slide 19 of 36
Slide 19 of 36
Lecture 5: Reinforcement Learning II - slide 20 of 36
Slide 20 of 36
Lecture 5: Reinforcement Learning II - slide 21 of 36
Slide 21 of 36
Lecture 5: Reinforcement Learning II - slide 22 of 36
Slide 22 of 36
Lecture 5: Reinforcement Learning II - slide 23 of 36
Slide 23 of 36
Lecture 5: Reinforcement Learning II - slide 24 of 36
Slide 24 of 36
Lecture 5: Reinforcement Learning II - slide 25 of 36
Slide 25 of 36
Lecture 5: Reinforcement Learning II - slide 26 of 36
Slide 26 of 36
Lecture 5: Reinforcement Learning II - slide 27 of 36
Slide 27 of 36
Lecture 5: Reinforcement Learning II - slide 28 of 36
Slide 28 of 36
Lecture 5: Reinforcement Learning II - slide 29 of 36
Slide 29 of 36
Lecture 5: Reinforcement Learning II - slide 30 of 36
Slide 30 of 36
Lecture 5: Reinforcement Learning II - slide 31 of 36
Slide 31 of 36
Lecture 5: Reinforcement Learning II - slide 32 of 36
Slide 32 of 36
Lecture 5: Reinforcement Learning II - slide 33 of 36
Slide 33 of 36
Lecture 5: Reinforcement Learning II - slide 34 of 36
Slide 34 of 36
Lecture 5: Reinforcement Learning II - slide 35 of 36
Slide 35 of 36
Lecture 5: Reinforcement Learning II - slide 36 of 36
Slide 36 of 36

Outcomes

Actor critic learning loop

Algorithm map

Method Policy/data regime Practical signature
REINFORCE On-policy Simple, high-variance returns
PPO On-policy actor-critic Clipped policy-ratio updates
SAC Off-policy actor-critic Replay, twin critics, entropy
Offline RL Fixed logged data Conservative or behavior-aware objectives

Core notes

Policy-gradient methods optimize a parameterized policy directly. The score-function estimator weights the gradient of an action’s log probability by its return. Subtracting a baseline leaves the estimator unbiased while reducing variance; actor-critic methods learn a value function for this role.

PPO is an on-policy method that limits how far the policy moves on a batch of recent experience, usually through a clipped probability-ratio objective. It is comparatively straightforward but discards data quickly. SAC is off-policy: it reuses replay data, learns critics, and adds entropy to encourage diverse behavior. It is often attractive for continuous control but is sensitive to critic errors and data quality.

Offline RL learns from a fixed dataset. The main danger is evaluating actions not supported by that dataset: an inaccurate critic may assign them unrealistically high value. Conservative objectives, behavior constraints, uncertainty, and careful dataset coverage help. Human intervention can supply corrective data near failures, but the intervention policy and safety boundary become part of the system.

Policy gradient and advantage

∇θJ(θ) = E[Σₜ ∇θ log πθ(aₜ│sₜ) Aₜ]

The advantage Aₜ asks whether an action was better than expected at that state. Generalized advantage estimation trades bias against variance using a decay parameter λ. Normalize advantages per batch only with a clear implementation convention.

PPO compares the new policy with the behavior policy that collected the batch:

rₜ(θ) = πθ(aₜ│sₜ) / π_old(aₜ│sₜ)

Its clipped objective discourages an update from changing action probability too far in one step. Track approximate KL divergence, clip fraction, policy entropy, value loss, and explained variance; reward alone cannot reveal a broken critic.

SAC optimizes return plus entropy. Twin critics reduce positive bias, replay improves reuse, and a learned temperature can target an entropy level. Check action squashing and log-probability corrections carefully for bounded continuous actions.

PPO versus SAC

Concept checks

  1. Why does reusing an old batch violate the simplest on-policy gradient derivation?
  2. When can entropy improve robustness, and when can it harm a precision task?

Build milestone

Train PPO or SAC on the same task contract used earlier. Compare it with the non-learning controller and behavior-cloning policy using equal evaluation episodes. Report environment steps and wall-clock time.

Ablate one stabilization choice: advantage normalization, entropy bonus, replay ratio, or target-network update rate. Explain the mechanism before interpreting the curve.


← Reinforcement Learning I · Next: Generative Models →