LLMs for Generalizable Language-Conditioned Policy Learning under Minimal Data Requirements

Pouplin, Thomas, Kobalczyk, Katarzyna, Sun, Hao, van der Schaar, Mihaela

Dec-9-2024–arXiv.org Artificial Intelligence

To develop autonomous agents capable of executing complex, multi-step decision-making tasks as specified by humans in natural language, existing reinforcement learning approaches typically require expensive labeled datasets or access to real-time experimentation. Moreover, conventional methods often face difficulties in generalizing to unseen goals and states, thereby limiting their practical applicability. This paper presents TEDUO, a novel training pipeline for offline language-conditioned policy learning. TEDUO operates on easy-to-obtain, unlabeled datasets and is suited for the so-called in-the-wild evaluation, wherein the agent encounters previously unseen goals and states. To address the challenges posed by such data and evaluation settings, our method leverages the prior knowledge and instruction-following capabilities of large language models (LLMs) to enhance the fidelity of pre-collected offline data and enable flexible generalization to new goals and states. Empirical results demonstrate that the dual role of LLMs in our framework-as data enhancers and generalizers-facilitates both effective and data-efficient learning of generalizable language-conditioned policies.

large language model, machine learning, reinforcement learning, (17 more...)

arXiv.org Artificial Intelligence

Dec-9-2024

arXiv.org PDF

Add feedback

Country:
- North America
  - United States
    - Illinois > Cook County
      - Chicago (0.04)
    - Florida > Broward County
      - Fort Lauderdale (0.04)
  - Canada > Alberta
    - Census Division No. 11 > Edmonton Metropolitan Region > Edmonton (0.04)
- Europe > United Kingdom
  - England > Cambridgeshire > Cambridge (0.14)
- Asia
  - Singapore (0.04)
  - Indonesia > Bali (0.04)

Genre:
- Research Report > New Finding (0.65)

Industry:
- Education (0.92)

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language > Large Language Model (1.00)
  - Machine Learning
    - Reinforcement Learning (1.00)
    - Neural Networks > Deep Learning (1.00)