Skip to the content.

Week 02 — Robot Control and MDPs

← Introduction · Week 2 of 11 · Next: Imitation Learning →

Complete lecture deck

Complete deck: 46 slides rendered locally. Select any slide to open the full-resolution image.
Lecture 2: Robot Control and MDPs - slide 1 of 46
Slide 1 of 46
Lecture 2: Robot Control and MDPs - slide 2 of 46
Slide 2 of 46
Lecture 2: Robot Control and MDPs - slide 3 of 46
Slide 3 of 46
Lecture 2: Robot Control and MDPs - slide 4 of 46
Slide 4 of 46
Lecture 2: Robot Control and MDPs - slide 5 of 46
Slide 5 of 46
Lecture 2: Robot Control and MDPs - slide 6 of 46
Slide 6 of 46
Lecture 2: Robot Control and MDPs - slide 7 of 46
Slide 7 of 46
Lecture 2: Robot Control and MDPs - slide 8 of 46
Slide 8 of 46
Lecture 2: Robot Control and MDPs - slide 9 of 46
Slide 9 of 46
Lecture 2: Robot Control and MDPs - slide 10 of 46
Slide 10 of 46
Lecture 2: Robot Control and MDPs - slide 11 of 46
Slide 11 of 46
Lecture 2: Robot Control and MDPs - slide 12 of 46
Slide 12 of 46
Lecture 2: Robot Control and MDPs - slide 13 of 46
Slide 13 of 46
Lecture 2: Robot Control and MDPs - slide 14 of 46
Slide 14 of 46
Lecture 2: Robot Control and MDPs - slide 15 of 46
Slide 15 of 46
Lecture 2: Robot Control and MDPs - slide 16 of 46
Slide 16 of 46
Lecture 2: Robot Control and MDPs - slide 17 of 46
Slide 17 of 46
Lecture 2: Robot Control and MDPs - slide 18 of 46
Slide 18 of 46
Lecture 2: Robot Control and MDPs - slide 19 of 46
Slide 19 of 46
Lecture 2: Robot Control and MDPs - slide 20 of 46
Slide 20 of 46
Lecture 2: Robot Control and MDPs - slide 21 of 46
Slide 21 of 46
Lecture 2: Robot Control and MDPs - slide 22 of 46
Slide 22 of 46
Lecture 2: Robot Control and MDPs - slide 23 of 46
Slide 23 of 46
Lecture 2: Robot Control and MDPs - slide 24 of 46
Slide 24 of 46
Lecture 2: Robot Control and MDPs - slide 25 of 46
Slide 25 of 46
Lecture 2: Robot Control and MDPs - slide 26 of 46
Slide 26 of 46
Lecture 2: Robot Control and MDPs - slide 27 of 46
Slide 27 of 46
Lecture 2: Robot Control and MDPs - slide 28 of 46
Slide 28 of 46
Lecture 2: Robot Control and MDPs - slide 29 of 46
Slide 29 of 46
Lecture 2: Robot Control and MDPs - slide 30 of 46
Slide 30 of 46
Lecture 2: Robot Control and MDPs - slide 31 of 46
Slide 31 of 46
Lecture 2: Robot Control and MDPs - slide 32 of 46
Slide 32 of 46
Lecture 2: Robot Control and MDPs - slide 33 of 46
Slide 33 of 46
Lecture 2: Robot Control and MDPs - slide 34 of 46
Slide 34 of 46
Lecture 2: Robot Control and MDPs - slide 35 of 46
Slide 35 of 46
Lecture 2: Robot Control and MDPs - slide 36 of 46
Slide 36 of 46
Lecture 2: Robot Control and MDPs - slide 37 of 46
Slide 37 of 46
Lecture 2: Robot Control and MDPs - slide 38 of 46
Slide 38 of 46
Lecture 2: Robot Control and MDPs - slide 39 of 46
Slide 39 of 46
Lecture 2: Robot Control and MDPs - slide 40 of 46
Slide 40 of 46
Lecture 2: Robot Control and MDPs - slide 41 of 46
Slide 41 of 46
Lecture 2: Robot Control and MDPs - slide 42 of 46
Slide 42 of 46
Lecture 2: Robot Control and MDPs - slide 43 of 46
Slide 43 of 46
Lecture 2: Robot Control and MDPs - slide 44 of 46
Slide 44 of 46
Lecture 2: Robot Control and MDPs - slide 45 of 46
Slide 45 of 46
Lecture 2: Robot Control and MDPs - slide 46 of 46
Slide 46 of 46

Outcomes

Markov decision process interaction

Quick reference

Symbol Meaning
sₜ, aₜ State and action at time t
P(s′│s,a) Transition distribution
R(s,a,s′) Immediate reward
γ Discount factor in [0,1)
π(a│s) Policy
Vπ(s), Qπ(s,a) State and action value under π

Core notes

Classical feedback control chooses actions from the current error. A proportional controller reacts to present error; integral action corrects persistent bias; derivative action damps fast change. This is often the strongest first baseline for a low-level robot task.

An MDP is described by states, actions, transitions, rewards, and a discount or finite horizon. A policy induces trajectories; the return aggregates future rewards; value functions predict expected return. Bellman equations express a value as immediate reward plus the value of the next state. Dynamic programming repeatedly applies this structure through policy evaluation, policy iteration, or value iteration.

The Markov assumption is a modeling decision, not a property guaranteed by the environment. A single camera frame may hide velocity, contact state, or intent. Frame stacking, recurrence, state estimation, or belief-state methods can repair missing information.

Rewards are specifications and can be exploited. Keep the true evaluation metric outside the shaped training reward, test simple pathological policies, and log individual reward terms.

Key relations

Return: Gₜ = rₜ + γrₜ₊₁ + γ²rₜ₊₂ + …

Bellman optimality: V*(s) = maxₐ E[r + γV*(s′) │ s,a]

Value iteration alternates one operation: replace every state’s value with the best one-step lookahead using the previous value table. Stop when the largest update is below a tolerance, then choose the maximizing action in each state.

Control baseline

For error e(t) = target - measurement, a PID command combines present, accumulated, and changing error:

u(t) = Kₚe(t) + Kᵢ∫e(τ)dτ + K_d de(t)/dt

Tune with actuator limits and sampling time in the loop. Add anti-windup for saturated integral action and filter noisy derivative estimates. A learned policy should beat this baseline on a reason the task actually values—robustness, adaptation, or a harder observation interface—not merely because the PID was untuned.

Common modeling mistakes

Paper discussion

Use random search for RL and Deep RL Doesn’t Work Yet to ask: how strong must a baseline be before a learned controller is convincing?

Build milestone

Implement a PID controller for one continuous task and value iteration for one tiny discrete MDP. Plot tracking error for the controller and verify the Bellman residual for the MDP solution.

Report rise time, steady-state error, overshoot, return, success rate, and the final Bellman residual. Perturb one physical parameter to test robustness.


← Introduction · Next: Imitation Learning →