Goto

Collaborating Authors

 Reinforcement Learning


Learning with Temporal Derivatives in Pulse-Coded Neuronal Systems

Neural Information Processing Systems

Reifsnider A number of learning models have recently been proposed which involve calculations of temporal differences (or derivatives in continuous-time models).



Stochastic Learning Networks and their Electronic Implementation

Neural Information Processing Systems

This paper focuses on the issue of learning in these networks especially with regard to their implementation in an electronic system. Learning phenomena that have been studied include associative memoryllJ.


Stochastic Learning Networks and their Electronic Implementation

Neural Information Processing Systems

This paper focuses on the issue of learning in these networks especially with regard to their implementation in an electronic system. Learning phenomena that have been studied include associative memoryllJ.


Stochastic Learning Networks and their Electronic Implementation

Neural Information Processing Systems

This paper focuses on the issue of learning in these networks especially with regard to their implementation in an electronic system. Learning phenomena that have been studied include associative memoryllJ.



Learning to predict by the methods of temporal difference

Classics

This article introduces a class of incremental learning procedures specializedfor prediction that is, for using past experience with an incompletely knownsystem to predict its future behavior. Whereas conventional prediction-learningmethods assign credit by means of the difference between predicted and actual outcomes,tile new methods assign credit by means of the difference between temporallysuccessive predictions. Although such temporal-difference method~ have been used inSamuel's checker player, Holland's bucket brigade, and the author's Adaptive HeuristicCritic, they have remained poorly understood. Here we prove their convergenceand optimality for special cases and relate them to supervised-learning methods. Formost real-world prediction problems, telnporal-differenee methods require less memoryand less peak computation than conventional methods and they produce moreaccurate predictions. We argue that most problems to which supervised learningis currently applied are really prediction problemsMachine Learning 3: 9-44, erratum p. 377