Provably Efficient Neural GTD for Off-Policy Learning

Oct-10-2024, 13:46:46 GMT–Neural Information Processing Systems

This paper studies a gradient temporal difference (GTD) algorithm using neural network (NN) function approximators to minimize the mean squared Bellman error (MSBE). For off-policy learning, we show that the minimum MSBE problem can be recast into a min-max optimization involving a pair of over-parameterized primal-dual NNs. The resultant formulation can then be tackled using a neural GTD algorithm. We analyze the convergence of the proposed algorithm with a 2-layer ReLU NN architecture using m neurons and prove that it computes an approximate optimal solution to the minimum MSBE problem as m \rightarrow \infty .

algorithm, off-policy learning, provably efficient neural gtd, (1 more...)

Neural Information Processing Systems

Oct-10-2024, 13:46:46 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.32)