Goto

Collaborating Authors

 Europe


Enhancing Q-Learning for Optimal Asset Allocation

Neural Information Processing Systems

This paper enhances the Q-Iearning algorithm for optimal asset allocation proposedin (Neuneier, 1996 [6]). The new formulation simplifies the approach by using only one value-function for many assets and allows model-freepolicy-iteration. After testing the new algorithm on real data, the possibility of risk management within the framework of Markov decision problems is analyzed. The proposed methods allows the construction of a multi-period portfolio management system which takes into account transaction costs, the risk preferences of the investor, and several constraints on the allocation. 1 Introduction


Reinforcement Learning for Call Admission Control and Routing in Integrated Service Networks

Neural Information Processing Systems

Peter Dayan E25-210, MIT Cambridge, MA 02139 We provide a model of the standard watermaze task, and of a more challenging task involving novel platform locations, in which rats exhibit one-trial learning after a few days of training. The model uses hippocampal place cells to support reinforcement learning, and also, in an integrated manner, to build and use allocentric coordinates. 1 INTRODUCTION


Asymptotic Theory for Regularization: One-Dimensional Linear Case

Neural Information Processing Systems

The generalization ability of a neural network can sometimes be improved dramatically by regularization. To analyze the improvement oneneeds more refined results than the asymptotic distribution ofthe weight vector. Here we study the simple case of one-dimensional linear regression under quadratic regularization, i.e., ridge regression. We study the random design, misspecified case, where we derive expansions for the optimal regularization parameter andthe ensuing improvement. It is possible to construct examples where it is best to use no regularization.


Experiences with Bayesian Learning in a Real World Application

Neural Information Processing Systems

Sleep staging is usually based on rules defined by Rechtschaffen and Kales (see [8]). Rechtschaffen and Kales rules define 4 sleep stages, stage one to four, as well as rapid eye movement (REM) and wakefulness. In [1] J. Bentrup and S. Ray report that every year nearly one million US citizens consulted their physicians concerning their sleep. Since sleep staging is a tedious task (one all night recording on average takes abou t 3 hours to score manually), much effort was spent in designing automatic sleep stagers. Sleep staging is a classification problem which was solved using classical statistical t.echniques or techniques emerged from the field of artificial intelligence (AI) . Among classical techniques especially the k nearest neighbor technique was used. In [1] J. Bentrup and S. Ray report that the classical technique outperformed their AI approaches. Among techniques from the field of AI, researchers used inductive learning to build tree based classifiers (e.g.


Unsupervised On-line Learning of Decision Trees for Hierarchical Data Analysis

Neural Information Processing Systems

An adaptive online algorithm is proposed to estimate hierarchical data structures for non-stationary data sources. The approach is based on the principle of minimum cross entropy to derive a decision tree for data clustering and it employs a metalearning idea (learning to learn) to adapt to changes in data characteristics. Its efficiency is demonstrated by grouping non-stationary artifical data and by hierarchical segmentation of LANDSAT images. 1 Introduction Unsupervised learning addresses the problem to detect structure inherent in unlabeled andunclassified data. N. The encoding usually is represented by an assignment matrix M (Mia), where Mia 1 if and only if Xi belongs to cluster L: 1 MiaV (Xi, Ya) measures the quality of a data partition, Le., optimal assignments and prototypes (M,y)OPt argminM,y1i (M,Y) minimize the inhomogeneity of clusters w.r.t. a given distance measure V. For reasons of simplicity we restrict the presentation to the ' sum-of-squared-error criterion V(x, y) To facilitate this minimization a deterministic annealing approach was proposed in [5] which maps the discrete optimization problem, i.e. how to determine the data assignments, viathe Maximum Entropy Principle [2] to a continuous parameter es- Unsupervised Online Learning ofDecision Trees for Data Analysis 515 timation problem.



Statistical Models of Conditioning

Neural Information Processing Systems

Conditioning experiments probe the ways that animals make predictions aboutrewards and punishments and use those predictions to control their behavior. One standard model of conditioning paradigms which involve many conditioned stimuli suggests that individual predictions should be added together. Various key results show that this model fails in some circumstances, and motivate analternative model, in which there is attentional selection between different available stimuli. The new model is a form of mixture of experts, has a close relationship with some other existing psychologicalsuggestions, and is statistically well-founded.


Effects of Spike Timing Underlying Binocular Integration and Rivalry in a Neural Model of Early Visual Cortex

Neural Information Processing Systems

In normal vision, the inputs from the two eyes are integrated intoa single percept. When dissimilar images are presented to the two eyes, however, perceptual integration givesway to alternation between monocular inputs, a phenomenon called binocular rivalry. Although recent evidence indicates that binocular rivalry involves a modulation ofneuronal responses in extrastriate cortex, the basic mechanisms responsible for differential processing of con:6.icting



Report on the Seventh International Workshop on Nonmonotonic Reasoning

AI Magazine

Fourth, causality is still an important issue; some formal models of causality have surprisingly close connections to standard nonmonotonic techniques. Fifth, the nonmonotonic logics being used most widely are the classical ones: default logic, circumscription, and by Isaac Levi; (3) Nonmonotonic Reasoning autoepistemic logic. Maybe the most remarkable trend he Seventh International Workshop was held in Trento, Italy, Tolerance by John McCarthy; (4) that became apparent during the on 30 May to 1 June 1998 in conjunction Learning to Make Nonmonotonic workshop was the new excitement with the Sixth International Inferences by Dan Roth; and (5) From among the participants. The depression Conference on the Principles of Features and Fluents to Thinking that plagued a number of people Knowledge Representation and Reasoning When Flying--Reasoning about in the field seems to be over. The workshop was Actions in an Intelligent UAV by Erik common feeling was that the theory sponsored by the American Association Sandewall.