Updates on Policy Gradients
I've been swamped with a bit of a travel binge and am hopelessly behind on blogging. After my last post on nominal control, I received an email from Pavel Christof pointing out that if we switch from stochastic gradient descent to Adam, policy gradient works much better. Indeed, I implemented this myself, and he's totally right. Let's revisit the last post with a revised Jupyter notebook. First, I coded up Adam in pure python to avoid introducing any deep learning package dependencies (it's only 4 lines of python, after all).
Mar-17-2018, 15:57:17 GMT
- Technology: