Learning values across many orders of magnitude
Hado P. van Hasselt, Arthur Guez, Arthur Guez, Matteo Hessel, Volodymyr Mnih, David Silver
–Neural Information Processing Systems
Most learning algorithms are not invariant to the scale of the signal that is being approximated. We propose to adaptively normalize the targets used in the learning updates. This is important in value-based reinforcement learning, where the magnitude of appropriate value approximations can change over time when we update the policy of behavior.
Neural Information Processing Systems
Nov-21-2025, 06:17:09 GMT
- Country:
- Europe
- Spain > Catalonia
- Barcelona Province > Barcelona (0.04)
- United Kingdom > England
- Cambridgeshire > Cambridge (0.14)
- Spain > Catalonia
- North America
- Puerto Rico (0.04)
- United States (0.04)
- Europe
- Industry:
- Leisure & Entertainment > Games > Computer Games (0.47)
- Technology: