Skip to the content.

Week 03 — Imitation Learning

← Control and MDPs · Week 3 of 11 · Next: Reinforcement Learning I →

Complete lecture deck

Complete deck: 45 slides rendered locally. Select any slide to open the full-resolution image.
Lecture 3: Imitation Learning - slide 1 of 45
Slide 1 of 45
Lecture 3: Imitation Learning - slide 2 of 45
Slide 2 of 45
Lecture 3: Imitation Learning - slide 3 of 45
Slide 3 of 45
Lecture 3: Imitation Learning - slide 4 of 45
Slide 4 of 45
Lecture 3: Imitation Learning - slide 5 of 45
Slide 5 of 45
Lecture 3: Imitation Learning - slide 6 of 45
Slide 6 of 45
Lecture 3: Imitation Learning - slide 7 of 45
Slide 7 of 45
Lecture 3: Imitation Learning - slide 8 of 45
Slide 8 of 45
Lecture 3: Imitation Learning - slide 9 of 45
Slide 9 of 45
Lecture 3: Imitation Learning - slide 10 of 45
Slide 10 of 45
Lecture 3: Imitation Learning - slide 11 of 45
Slide 11 of 45
Lecture 3: Imitation Learning - slide 12 of 45
Slide 12 of 45
Lecture 3: Imitation Learning - slide 13 of 45
Slide 13 of 45
Lecture 3: Imitation Learning - slide 14 of 45
Slide 14 of 45
Lecture 3: Imitation Learning - slide 15 of 45
Slide 15 of 45
Lecture 3: Imitation Learning - slide 16 of 45
Slide 16 of 45
Lecture 3: Imitation Learning - slide 17 of 45
Slide 17 of 45
Lecture 3: Imitation Learning - slide 18 of 45
Slide 18 of 45
Lecture 3: Imitation Learning - slide 19 of 45
Slide 19 of 45
Lecture 3: Imitation Learning - slide 20 of 45
Slide 20 of 45
Lecture 3: Imitation Learning - slide 21 of 45
Slide 21 of 45
Lecture 3: Imitation Learning - slide 22 of 45
Slide 22 of 45
Lecture 3: Imitation Learning - slide 23 of 45
Slide 23 of 45
Lecture 3: Imitation Learning - slide 24 of 45
Slide 24 of 45
Lecture 3: Imitation Learning - slide 25 of 45
Slide 25 of 45
Lecture 3: Imitation Learning - slide 26 of 45
Slide 26 of 45
Lecture 3: Imitation Learning - slide 27 of 45
Slide 27 of 45
Lecture 3: Imitation Learning - slide 28 of 45
Slide 28 of 45
Lecture 3: Imitation Learning - slide 29 of 45
Slide 29 of 45
Lecture 3: Imitation Learning - slide 30 of 45
Slide 30 of 45
Lecture 3: Imitation Learning - slide 31 of 45
Slide 31 of 45
Lecture 3: Imitation Learning - slide 32 of 45
Slide 32 of 45
Lecture 3: Imitation Learning - slide 33 of 45
Slide 33 of 45
Lecture 3: Imitation Learning - slide 34 of 45
Slide 34 of 45
Lecture 3: Imitation Learning - slide 35 of 45
Slide 35 of 45
Lecture 3: Imitation Learning - slide 36 of 45
Slide 36 of 45
Lecture 3: Imitation Learning - slide 37 of 45
Slide 37 of 45
Lecture 3: Imitation Learning - slide 38 of 45
Slide 38 of 45
Lecture 3: Imitation Learning - slide 39 of 45
Slide 39 of 45
Lecture 3: Imitation Learning - slide 40 of 45
Slide 40 of 45
Lecture 3: Imitation Learning - slide 41 of 45
Slide 41 of 45
Lecture 3: Imitation Learning - slide 42 of 45
Slide 42 of 45
Lecture 3: Imitation Learning - slide 43 of 45
Slide 43 of 45
Lecture 3: Imitation Learning - slide 44 of 45
Slide 44 of 45
Lecture 3: Imitation Learning - slide 45 of 45
Slide 45 of 45

Outcomes

Behavior cloning and DAgger comparison

Quick reference

Method Data collection Strength Main limitation
Behavior cloning Expert trajectories once Simple and stable Learner-state shift
DAgger Iterative learner rollouts + expert relabeling Recovery-state coverage Requires repeated expert access
Action chunking Demonstrations with sequence targets Temporal coherence Slower feedback for long chunks
Generative policy Demonstrations with multimodal targets Represents alternatives More complex inference

Core notes

Behavior cloning fits a policy to expert observation-action pairs. It is simple and stable, but the learned policy changes which states it visits. Small errors move the robot away from the demonstration distribution; unfamiliar states cause more errors, and mistakes compound over the horizon.

DAgger addresses this by rolling out the learner, querying the expert on states the learner actually visits, aggregating those labels, and retraining. It trades expert effort for better coverage of recovery states. When online relabeling is impossible, dataset diversity, perturbation/recovery demonstrations, conservative deployment, and uncertainty-aware fallback become important.

Robotic behavior is often multimodal: several actions can be correct in the same observation. A mean-squared-error policy may average them into an invalid action. Discrete bins, mixture models, energy-based models, diffusion policies, or action chunks can represent alternatives. Chunking also reduces the effective decision horizon, but long chunks reduce feedback frequency.

Causal confusion appears when a policy uses a correlate that predicts expert actions in the dataset but does not cause success. Evaluate under interventions: change backgrounds, object arrangements, demonstrator artifacts, or history while holding the task fixed.

Objectives and compounding error

Behavior cloning minimizes supervised negative log-likelihood:

L_BC(θ) = - E_(o,a)~D [log πθ(a│o)]

This expectation is over the expert dataset, while deployment observations come from the learner-induced distribution. A small per-step mistake probability can therefore produce a much larger trajectory-level failure rate over a long horizon.

DAgger closes the loop:

  1. Train on the current aggregated dataset.
  2. Roll out the learner, optionally mixed with the expert for safety.
  3. Ask the expert what action should have been taken at visited states.
  4. Add those pairs to the dataset and repeat.

Dataset and architecture decisions

Failure modes

Paper discussion

Contrast the diagnosis in Causal Confusion in Imitation Learning with the empirical case for pretrained visual representations in Pari et al..

Build milestone

Collect or synthesize demonstrations, train behavior cloning, then introduce initial-state noise. Add either DAgger or recovery data and compare closed-loop success—not just validation loss—over three seeds.

Plot performance against perturbation magnitude and dataset size. Include the expert, a random policy, and behavior cloning before claiming the interactive method helps.


← Control and MDPs · Next: Reinforcement Learning I →