Loading
russtedrake
This is a live streamed presentation. You will automatically follow the presenter and see the slide they're currently on.
MIT 6.881: Robotic Manipulation
Fall 2020, Lecture 14
Follow live at https://slides.com/russtedrake/fall20-lec14/live
(or later at https://slides.com/russtedrake/fall20-lec14)
A useful first step...
SE(3) pose is difficult to generalize across a category
So how do we even specify the task?
What's the cost function?
(Images of mugs on the rack?)
H. Wang, S. Sridhar, J. Huang, J. Valentin, S. Song, and L. J. Guibas, “Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, Jun. 2019, pp. 2637–2646, doi: 10.1109/CVPR.2019.00275.
Lucas Manuelli*, Wei Gao*, Peter R. Florence and Russ Tedrake. kPAM: KeyPoint Affordances for Category Level Manipulation. ISRR 2019
Manipulate potentially unknown rigid objects from a category (e.g. mugs, shoes) into desired target configurations.
https://keypointnet.github.io/
https://nanonets.com/blog/human-pose-estimation-2d-guide/
3D Keypoints provide rich, class-general semantics
Constraints & Cost on Keypoints
... and robust performance in practice
Lucas Manuelli*, Wei Gao*, Peter R. Florence and Russ Tedrake. kPAM: KeyPoint Affordances for Category Level Manipulation. ISRR 2019
subject to:
Keypoint "semantics" + dense 3D geometry
No template model nor pose appears in this pipeline.
Custom annotation
tool
RGBD image w/ instance segmentation
Grasp
Planner
3D Keypoint Detection Network
Image
3D
Inverse Kinematics Planner
Architecture based on Sun, Xiao, et al. "Integral human pose regression." ECCV, 2018
Custom annotation tool
Sample of results
(shoes on rack)
| # train objects | 10 |
| # test objects | 20 |
| # trials | 100 |
| placed on shelf | 98% |
| heel error (cm) | 1.09 ± (1.29) |
| toe error (cm) | 4.34 ± (3.05) |
to include collision-avoidance constraints
So far, keypoints are geometric and semantic
(mug handle, front toe of shoe), but required human labels
If we forgo semantics, can we self-supervise?
Z. Qin, K. Fang, Y. Zhu, L. Fei-Fei, and S. Savarese, “KETO: Learning Keypoint Representations for Tool Manipulation,” in 2020 IEEE International Conference on Robotics and Automation (ICRA), May 2020, pp. 7278–7285
Core technology: dense correspondences
(built on Schmidt, Newcombe, Fox, RA-L 2017)
Peter R. Florence*, Lucas Manuelli*, and Russ Tedrake. Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation. CoRL, 2018.
dense 3D reconstruction
+ pixelwise contrastive loss
New loss function sharpens correspondences
Now 3D correspondences, trained with multiview
Dense descriptors as self-supervised keypoints
Correspondences alone are sufficient to specify some tasks
Peter R. Florence, Lucas Manuelli, and Russ Tedrake. Self-Supervised Correspondence in Visuomotor Policy Learning. RA-L, April 2020
Levine*, Finn*, Darrel, Abbeel, JMLR 2016
Estimate object/hand pose
(but hard for category-level)
from hand-coded policies in sim
and teleop on the real robot
Standard "behavior-cloning" objective + data augmentation
"push box"
"flip box"
Policy is a small LSTM network (~100 LSTMs)
Learn descriptor keypoint dynamics + trajectory MPC
Learn descriptor keypoint dynamics + trajectory MPC
My take-aways:
Lucas at his defense: "Perception doesn't feel like the bottleneck anymore; it fees like the bottleneck is control."
Deep models
f=ma
Wanted: State representations for "category-level" dexterous manipulation
State representations
H.J. Terry Suh and Russ Tedrake. The surprising effectiveness of linear models for visual foresight in object pile manipulation. To appear in Workshop on the Algorithmic Foundations of Robotics (WAFR), 2020
Target Set
The big question:
Find a single policy (pixels to torques) that is invariant to the number of carrot pieces.
Try images as the state representation.
A control-Lyapunov function in image coordinates
Requires a forward model...
Crazy: Try a linear model (for each discretized action)
Each row of is an image, the "receptive field" of
My take-aways:
Today:
System
Auto-regressive (eg. ARMAX)
Lagrangian mechanics,
Recurrent neural networks (e.g. LSTM), ...
Feed-forward networks (e.g. \(y_n\)= image)
input
output
State-space
The failings of our physics-based models are mostly due to the unreasonable burden of estimating the "Lagrangian state" and parameters.
For e.g. onions, laundry, peanut butter, ...
The failings of our deep models are mostly due to our inability to due efficient/reliable planning, control design and analysis.
I want the next Newton to come around and to work on onions, laundry, peanut butter...
This fall: my manipulation class at MIT will be online.