Self-Adapting Improvement Loops for Robotic Learning
Luo, Calvin, Zeng, Zilai, Jia, Mingxi, Du, Yilun, Sun, Chen
–arXiv.org Artificial Intelligence
Advancements in video generative modeling capabilities have directly led to their increased utilization as visual planners for robotic applications [1, 2, 3, 4]. The synthesized visual plan, in the form of video frames generated with text conditioning, can be translated into executable actions via inverse dynamics models (IDMs). While the IDMs are generally robust across tasks, the data on which the video generative models are trained can greatly impact downstream robotic performance and generalization. When explicitly optimized on in-domain examples of expert behavior, such visual planners are able to synthesize successful plans for solving demonstrated tasks in a robust manner. However, for arbitrary robotic settings, large-scale expert-quality datasets may not be readily available, and collection may be prohibitively expensive. A paucity of data scale can limit video models trained only on in-domain videos from exhibiting generalized planning capabilities for novel tasks. Integrating knowledge from large-scale datasets of text and video collected from the internet has facilitated improved generalization, even in the absence of abundant in-domain videos. Recent work, Adapt2Act [5], creates a powerful, generalizable, text-conditioned visual planner by combining a large-scale model pretrained on web-scale video data with a video model trained on a small set of in-domain demonstrations via score composition. At a high level, the adapted video model draws upon large-scale motion priors and powerful zero-shot text conditioning capabilities from the web-pretrained video model to facilitate generalization.
arXiv.org Artificial Intelligence
Jun-10-2025
- Country:
- North America > Mexico (0.28)
- Genre:
- Research Report (0.64)
- Technology:
- Information Technology > Artificial Intelligence
- Robots (1.00)
- Representation & Reasoning (1.00)
- Machine Learning (1.00)
- Natural Language > Large Language Model (0.88)
- Information Technology > Artificial Intelligence