Learning to Predict Without Looking Ahead: World Models Without Forward Prediction

Oct-9-2024, 14:24:47 GMT–Neural Information Processing Systems

Much of model-based reinforcement learning involves learning a model of an agent's world, and training an agent to leverage this model to perform a task more efficiently. While these models are demonstrably useful for agents, every naturally occurring model of the world of which we are aware---e.g., a brain---arose as the byproduct of competing evolutionary pressures for survival, not minimization of a supervised forward-predictive loss via gradient descent. That useful models can arise out of the messy and slow optimization process of evolution suggests that forward-predictive modeling can arise as a side-effect of optimization under the right circumstances. Crucially, this optimization process need not explicitly be a forward-predictive loss. In this work, we introduce a modification to traditional reinforcement learning which we call observational dropout, whereby we limit the agents ability to observe the real environment at each timestep.

agent, forward prediction, world model, (4 more...)

Neural Information Processing Systems

Oct-9-2024, 14:24:47 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence
  - Cognitive Science > Problem Solving (0.49)
  - Machine Learning > Reinforcement Learning (0.57)