Skip to the content.

Week 04 — Reinforcement Learning I

← Imitation Learning · Week 4 of 11 · Next: Reinforcement Learning II →

Complete lecture deck

Complete deck: 45 slides rendered locally. Select any slide to open the full-resolution image.
Lecture 4: Reinforcement Learning I - slide 1 of 45
Slide 1 of 45
Lecture 4: Reinforcement Learning I - slide 2 of 45
Slide 2 of 45
Lecture 4: Reinforcement Learning I - slide 3 of 45
Slide 3 of 45
Lecture 4: Reinforcement Learning I - slide 4 of 45
Slide 4 of 45
Lecture 4: Reinforcement Learning I - slide 5 of 45
Slide 5 of 45
Lecture 4: Reinforcement Learning I - slide 6 of 45
Slide 6 of 45
Lecture 4: Reinforcement Learning I - slide 7 of 45
Slide 7 of 45
Lecture 4: Reinforcement Learning I - slide 8 of 45
Slide 8 of 45
Lecture 4: Reinforcement Learning I - slide 9 of 45
Slide 9 of 45
Lecture 4: Reinforcement Learning I - slide 10 of 45
Slide 10 of 45
Lecture 4: Reinforcement Learning I - slide 11 of 45
Slide 11 of 45
Lecture 4: Reinforcement Learning I - slide 12 of 45
Slide 12 of 45
Lecture 4: Reinforcement Learning I - slide 13 of 45
Slide 13 of 45
Lecture 4: Reinforcement Learning I - slide 14 of 45
Slide 14 of 45
Lecture 4: Reinforcement Learning I - slide 15 of 45
Slide 15 of 45
Lecture 4: Reinforcement Learning I - slide 16 of 45
Slide 16 of 45
Lecture 4: Reinforcement Learning I - slide 17 of 45
Slide 17 of 45
Lecture 4: Reinforcement Learning I - slide 18 of 45
Slide 18 of 45
Lecture 4: Reinforcement Learning I - slide 19 of 45
Slide 19 of 45
Lecture 4: Reinforcement Learning I - slide 20 of 45
Slide 20 of 45
Lecture 4: Reinforcement Learning I - slide 21 of 45
Slide 21 of 45
Lecture 4: Reinforcement Learning I - slide 22 of 45
Slide 22 of 45
Lecture 4: Reinforcement Learning I - slide 23 of 45
Slide 23 of 45
Lecture 4: Reinforcement Learning I - slide 24 of 45
Slide 24 of 45
Lecture 4: Reinforcement Learning I - slide 25 of 45
Slide 25 of 45
Lecture 4: Reinforcement Learning I - slide 26 of 45
Slide 26 of 45
Lecture 4: Reinforcement Learning I - slide 27 of 45
Slide 27 of 45
Lecture 4: Reinforcement Learning I - slide 28 of 45
Slide 28 of 45
Lecture 4: Reinforcement Learning I - slide 29 of 45
Slide 29 of 45
Lecture 4: Reinforcement Learning I - slide 30 of 45
Slide 30 of 45
Lecture 4: Reinforcement Learning I - slide 31 of 45
Slide 31 of 45
Lecture 4: Reinforcement Learning I - slide 32 of 45
Slide 32 of 45
Lecture 4: Reinforcement Learning I - slide 33 of 45
Slide 33 of 45
Lecture 4: Reinforcement Learning I - slide 34 of 45
Slide 34 of 45
Lecture 4: Reinforcement Learning I - slide 35 of 45
Slide 35 of 45
Lecture 4: Reinforcement Learning I - slide 36 of 45
Slide 36 of 45
Lecture 4: Reinforcement Learning I - slide 37 of 45
Slide 37 of 45
Lecture 4: Reinforcement Learning I - slide 38 of 45
Slide 38 of 45
Lecture 4: Reinforcement Learning I - slide 39 of 45
Slide 39 of 45
Lecture 4: Reinforcement Learning I - slide 40 of 45
Slide 40 of 45
Lecture 4: Reinforcement Learning I - slide 41 of 45
Slide 41 of 45
Lecture 4: Reinforcement Learning I - slide 42 of 45
Slide 42 of 45
Lecture 4: Reinforcement Learning I - slide 43 of 45
Slide 43 of 45
Lecture 4: Reinforcement Learning I - slide 44 of 45
Slide 44 of 45
Lecture 4: Reinforcement Learning I - slide 45 of 45
Slide 45 of 45

Outcomes

Deep Q learning data flow

Value-learning ladder

Method Target Data regime
Monte Carlo Complete sampled return On-policy episodes
TD(0) r + γV(s′) Online bootstrapping
Q-learning r + γ maxₐ′ Q(s′,a′) Off-policy transitions
DQN Q-learning target with neural networks Replay + target network

Core notes

Reinforcement learning optimizes expected return from interaction. Monte Carlo targets use complete sampled returns and can have high variance. Temporal-difference methods bootstrap from an estimate of the next state’s value, reducing target horizon while introducing bias and moving-target instability.

Q-learning learns the value of taking an action and then acting optimally. With function approximation, DQN stabilizes training using replayed transitions and a slower target network. Replay improves data reuse and weakens temporal correlation; the target network stops every update from immediately changing its own target.

The central practical loop is: collect transitions, store them, sample a batch, construct a Bellman target, minimize prediction error, and periodically update the target network. Exploration must be stated explicitly; epsilon-greedy is a baseline, not a universal solution.

RL results are distributions. Report learning curves, final performance, seed count, environment steps, wall-clock cost, and evaluation without exploration noise. Compare against random, scripted, and classical-control baselines where applicable.

DQN objective

For a replayed transition (s,a,r,s′,done), construct

y = r + γ(1-done) maxₐ′ Qθ⁻(s′,a′)

L(θ) = E[(Qθ(s,a) - y)²]

The online parameters θ receive gradient updates. The target parameters θ⁻ are copied or slowly averaged from the online network. The done mask prevents bootstrapping beyond a true terminal state; time-limit truncation may need separate treatment.

Minimal training loop

  1. Select an action with an explicit exploration policy.
  2. Step the environment and store the transition.
  3. Sample a random replay batch after a warm-up period.
  4. Compute masked targets with the target network.
  5. Update the online Q network and periodically update the target.
  6. Evaluate in separate episodes without exploration noise.

Diagnostics

Paper discussion

Compare gradient-free evolution strategies with value-based learning. Which method is easier to parallelize, and which uses each transition more efficiently?

Build milestone

Solve a tabular task with value iteration, then train DQN on a small discrete-control task. Add an ablation removing either replay or the target network and explain the resulting curve.

Use the same evaluation seeds for every checkpoint, but keep them out of training. Report sample efficiency and wall-clock efficiency separately.


← Imitation Learning · Next: Reinforcement Learning II →