Back to Home
Stanford AI Lab··Papers & Tech

Reward Isn't Free: Supervising Robot Learning with Language and Video from the Web

中文摘要

AI robots face challenges in generalizability. Research focuses on enabling robots to learn and adapt to new environments.

English Summary

This work utilizes web-based language and video to supervise robot learning, helping home robots generalize their knowledge to perform diverse interactive tasks in novel environments.

Original Excerpt

This work was conducted as part of SAIL and CRFM. Deep learning has enabled improvements in the capabilities of robots on a range of problems such as grasping 1 and locomotion 2 in recent years. However, building the quintessential home robot that can perform a range of interactive tasks, from cooking to cleaning, in novel environments has remained elusive. While a number of hardware and software challenges remain, a necessary component is robots that can generalize their prior knowledge to new environments, tasks, and objects in a zero or few shot manner. For example, a home robot tasked with setting the dining table cannot afford lengthy re-training for every new dish, piece of cutlery, or dining room it may need to interact with. A natural way to enable such generalization in our robots is to train them on rich data sources that contain a wide range of different environments, tasks, and objects. Indeed, this recipe of massive, diverse datasets combined with scalable offline learning algorithms (e.g. self-supervised or cheaply supervised learning) has been the backbone of the many recent successes of foundation models 3 in NLP 456789 and vision 101112. Replicating these impressiv…