russtedrake PRO
Roboticist at MIT and TRI
MIT 6.421
Fall 2026, Lecture 2
Russ Tedrake
Walking into the building just before lecture 1 last week, I overheard:
Note: It did take 2.5 minutes to run..., so it's a harbinger, but not quite ready to ship...
“Pick up the red block from the table and place it inside the bowl” using Inspect Robots
omni model
text, images,
video, audio, ...
robot sensors
text, images,
video, audio, ...
robot actions
as you've seen in, e.g. MIT 6.390 (https://introml.mit.edu/)
image from old result:
OpenAI - Learning Dexterity
note: BC is one type of imitation learning
note: BC is one type of imitation learning
Why is RL dominant for locomotion, but imitation learning more dominant for manipulation?
(And how will that evolve over the coming years...?)
"Single-task" learning
Multitask behavior cloning
vision encoder
language encoder
action
decoder
robot joint encoder
vision encoder
action
decoder
robot joint encoder
dexterity
generality
large-scale transformer-based multitask learning??
Initial scaling laws look very promising
(but it's very hard to evaluate from social media posts; there is still a lot to do!)
Often very simple, e.g. 3 layer, 255 unit, multilayer perceptrons (MLPs)
Levine*, Finn*, Darrel, Abbeel, JMLR 2016
perception network
(often pre-trained)
policy network
other robot sensors
learned state representation
actions
x history
omni model
text, images,
video, audio, ...
robot sensors
text, images,
video, audio, ...
robot actions
example: GR00T N1 architecture
Big data
Big transfer
Small data
No transfer
robot teleop
(the "transfer learning bet")
Open-X
simulation rollouts
novel devices
from the LBM 1.0 paper
more data from this task
more data from other tasks
fine-tuning
pretraining
ClearKitchenCounter
Single task
LBM finetuned
By russtedrake
MIT Robotic Manipulation Fall 2026 http://manipulation.mit.edu