Stabilizing Value Function Approximation with the BFBP Algorithm

Wang, Xin, Dietterich, Thomas G.

Neural Information Processing Systems 

Our BFBP (Batch Fit to Best Paths) algorithm alternates between an exploration phase (during which trajectories are generated to try to find fragments of the optimal policy) and a function fitting phase (during which a function approximator is fit to the best known paths from start states to terminal states). An advantage of this approach is that batch value-function fitting is a global process, which allows it to address the tradeoffs in function approximation that cannot be handled by local, online algorithms.

Similar Docs  Excel Report  more

TitleSimilaritySource
None found