Skip to the content.

Curriculum guide

Goal

Move from the mechanics of robot learning to the ability to design, implement, and evaluate a small modern system. The sequence is cumulative: later sessions reuse the same task contract, dataset format, and evaluation habits established in weeks 1–3.

What is on each lecture page

The locally rendered decks retain the copyright and reuse terms of the source material. The additional notes and diagrams are independent study-group material.

Tracks

Use one lightweight task for several weeks so algorithm differences remain interpretable. Good choices are a simulated reaching task, planar pushing, or a small grid-world for the earliest sessions. Define:

Reading list by week

Week Suggested discussion papers
1 No required paper; agree on task and evaluation conventions.
2 Simple random search provides a competitive approach to RL; Deep RL Doesn’t Work Yet; Curiosity-driven Exploration
3 Causal Confusion in Imitation Learning; Representation Learning for Visual Imitation; Transporter Networks
4 Evolution Strategies as a Scalable Alternative to RL; Learning Synergies Between Pushing and Grasping; Human-in-the-loop RL for Dexterous Manipulation
5 End-to-End Training of Deep Visuomotor Policies; Eureka; Latent Plans for Task-Agnostic Offline RL
6 Diffuser; Implicit Behavioral Cloning; Steering Your Diffusion Policy with Latent Space RL
7 Decision Transformer; ALOHA; Humanoid Locomotion as Next Token Prediction
8 Universal Policies via Text-Guided Video Generation; Training Agents Inside Scalable World Models; World Action Models are Zero-shot Policies
9 Language Conditioned Imitation Learning; Gato; π*0.6
10 In-Context Imitation Learning; Voyager; Training Strategies for Efficient Embodied Reasoning
11 A Path Towards Autonomous Machine Intelligence; The Bitter Lesson; Intelligence without Representation

The official course page is the source of truth for its schedule and links.

Capstone checkpoint

By week 11, each team should have a one-page proposal containing a falsifiable question, baseline, intervention, dataset/task, primary metric, compute estimate, ablations, and a stop condition. Prefer a small clean result over an unbounded platform build.