Stabilizing Value Function Approximation with the BFBP Algorithm
Wang, Xin, Dietterich, Thomas G.
–Neural Information Processing Systems
Our BFBP (Batch Fit to Best Paths) algorithm alternates between an exploration phase (during which trajectories are generated to try to find fragments of the optimal policy) and a function fitting phase (during which a function approximator is fit to the best known paths from start states to terminal states). An advantage of this approach is that batch value-function fitting is a global process, which allows it to address the tradeoffs in function approximation that cannot be handled by local, online algorithms.
Neural Information Processing Systems
Dec-31-2002
- Country:
- North America > United States
- California > San Francisco County
- San Francisco (0.15)
- Oregon (0.14)
- California > San Francisco County
- North America > United States
- Technology: