Skip to the content.

Week 10 — Embodied Reasoning and Test-time Scaling

← Generalist Policies · Week 10 of 11 · Next: Frontiers →

Complete lecture deck

Complete deck: 57 slides rendered locally. Select any slide to open the full-resolution image.
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 1 of 57
Slide 1 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 2 of 57
Slide 2 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 3 of 57
Slide 3 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 4 of 57
Slide 4 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 5 of 57
Slide 5 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 6 of 57
Slide 6 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 7 of 57
Slide 7 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 8 of 57
Slide 8 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 9 of 57
Slide 9 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 10 of 57
Slide 10 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 11 of 57
Slide 11 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 12 of 57
Slide 12 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 13 of 57
Slide 13 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 14 of 57
Slide 14 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 15 of 57
Slide 15 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 16 of 57
Slide 16 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 17 of 57
Slide 17 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 18 of 57
Slide 18 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 19 of 57
Slide 19 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 20 of 57
Slide 20 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 21 of 57
Slide 21 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 22 of 57
Slide 22 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 23 of 57
Slide 23 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 24 of 57
Slide 24 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 25 of 57
Slide 25 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 26 of 57
Slide 26 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 27 of 57
Slide 27 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 28 of 57
Slide 28 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 29 of 57
Slide 29 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 30 of 57
Slide 30 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 31 of 57
Slide 31 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 32 of 57
Slide 32 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 33 of 57
Slide 33 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 34 of 57
Slide 34 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 35 of 57
Slide 35 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 36 of 57
Slide 36 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 37 of 57
Slide 37 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 38 of 57
Slide 38 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 39 of 57
Slide 39 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 40 of 57
Slide 40 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 41 of 57
Slide 41 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 42 of 57
Slide 42 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 43 of 57
Slide 43 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 44 of 57
Slide 44 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 45 of 57
Slide 45 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 46 of 57
Slide 46 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 47 of 57
Slide 47 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 48 of 57
Slide 48 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 49 of 57
Slide 49 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 50 of 57
Slide 50 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 51 of 57
Slide 51 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 52 of 57
Slide 52 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 53 of 57
Slide 53 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 54 of 57
Slide 54 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 55 of 57
Slide 55 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 56 of 57
Slide 56 of 57
Lecture 10: Embodied Reasoning and Test-time Scaling - slide 57 of 57
Slide 57 of 57

Outcomes

Embodied reasoning loop

Test-time compute options

Mechanism Added capability Typical risk
Multiple samples Alternative plans/actions Correlated candidates
Search Multi-step lookahead Branching and latency
Verifier/value model Rank candidates Reward hacking or miscalibration
Tools/simulator Ground facts and consequences Tool-model mismatch
Replanning Recover from execution mismatch Control delay

Core notes

Some robot tasks require more than a single reactive policy pass: long-horizon instructions, tool choice, recovery, spatial constraints, or unfamiliar combinations. Test-time scaling spends additional computation to sample plans, search a tree, critique candidates, use tools, simulate outcomes, or replan after new observations.

The useful pattern is propose → ground → verify → execute → observe → replan. A language model may propose subgoals, but grounded modules must connect them to visible objects, reachable poses, controller capabilities, and safety constraints. Verification can use a value model, learned world model, geometric checker, simulator, or explicit predicate tests.

More inference is not automatically better. Candidate samples may be correlated, verifiers may reward polished explanations instead of feasible actions, and latency may exceed the control budget. Compare against an equal-latency reactive baseline and report accuracy or success as a function of samples, search depth, and wall-clock time.

In-context imitation provides another route: condition on example trajectories and predict the next action without updating weights. The key test is whether the policy recombines task structure or merely matches surface similarity.

A grounded inference loop

  1. Parse the instruction into candidate subgoals and constraints.
  2. Ground referenced objects and relations in the current observation.
  3. Filter candidates through reachability, collision, and skill-availability checks.
  4. Score surviving plans with a verifier or world model.
  5. Execute the smallest useful step and collect fresh evidence.
  6. Replan when predicates fail, uncertainty rises, or the scene changes.

The verifier should consume grounded state—not only the model’s own explanation. Useful checks include geometric feasibility, resource preconditions, predicted task predicates, safety constraints, and calibrated uncertainty.

Scaling curves

For each test-time budget B, record:

utility(B) = task success(B) - λ · latency(B) - μ · safety cost(B)

Sweep candidate count or search depth and compare against a reactive baseline given the same latency. Stop scaling when marginal success no longer justifies compute or control delay.

Adversarial checks

Paper discussion

Compare embodied learning from examples in In-Context Imitation Learning with iterative skill acquisition in Voyager.

Build milestone

Add candidate generation and a simple verifier to a multi-step task. Sweep the number of candidates and plot success versus latency. Include at least one adversarial case where the verifier selects a bad plan.

Report verifier accuracy separately from end-to-end task success. Include the cost of perception, candidate generation, verification, and control.


← Generalist Policies · Next: Frontiers →