Goto

Collaborating Authors

 Statistical Learning


A Hodgkin-Huxley Type Neuron Model That Learns Slow Non-Spike Oscillation

Neural Information Processing Systems

A gradient descent algorithm for parameter estimation which is similar to those used for continuous-time recurrent neural networks was derived for Hodgkin-Huxley type neuron models. Using mem(cid:173) brane potential trajectories as targets, the parameters (maximal conductances, thresholds and slopes of activation curves, time con(cid:173) stants) were successfully estimated. The algorithm was applied to modeling slow non-spike oscillation of an identified neuron in the lobster stomatogastric ganglion. A model with three ionic currents was trained with experimental data. It revealed a novel role of A-current for slow oscillation below -50 mY.


Combined Neural Networks for Time Series Analysis

Neural Information Processing Systems

We propose a method for improving the performance of any net(cid:173) work designed to predict the next value of a time series. Vve advo(cid:173) cate analyzing the deviations of the network's predictions from the data in the training set. This can be carried out by a secondary net(cid:173) work trained on the time series of these residuals. The combined system of the two networks is viewed as the new predictor. We demonstrate the simplicity and success of this method, by apply(cid:173) ing it to the sunspots data.


Solvable Models of Artificial Neural Networks

Neural Information Processing Systems

Solvable models of nonlinear learning machines are proposed, and learning in artificial neural networks is studied based on the theory of ordinary differential equations. A learning algorithm is con(cid:173) structed, by which the optimal parameter can be found without any recursive procedure. The solvable models enable us to analyze the reason why experimental results by the error backpropagation often contradict the statistical learning theory.


Locally Adaptive Nearest Neighbor Algorithms

Neural Information Processing Systems

Four versions of a k-nearest neighbor algorithm with locally adap(cid:173) tive k are introduced and compared to the basic k-nearest neigh(cid:173) bor algorithm (kNN). Locally adaptive kNN algorithms choose the value of k that should be used to classify a query by consulting the results of cross-validation computations in the local neighborhood of the query. Local kNN methods are shown to perform similar to kNN in experiments with twelve commonly used data sets. Encour(cid:173) aging results in three constructed tasks show that local methods can significantly outperform kNN in specific applications. Local methods can be recommended for on-line learning and for appli(cid:173) cations where different regions of the input space are covered by patterns solving different sub-tasks.


Non-Linear Statistical Analysis and Self-Organizing Hebbian Networks

Neural Information Processing Systems

Neurons learning under an unsupervised Hebbian learning rule can perform a nonlinear generalization of principal component analysis. This relationship between nonlinear PCA and nonlinear neurons is reviewed. The stable fixed points of the neuron learning dynamics correspond to the maxima of the statist,ic optimized under non(cid:173) linear PCA. However, in order to predict. This is shown for a simple model. Methods of statistical mechanics can be used to find the optima of the objective function of non-linear PCA.


GDS: Gradient Descent Generation of Symbolic Classification Rules

Neural Information Processing Systems

Imagine you have designed a neural network that successfully learns a complex classification task. What are the relevant input features the classifier relies on and how are these features combined to pro(cid:173) duce the classification decisions? There are applications where a deeper insight into the structure of an adaptive system and thus into the underlying classification problem may well be as important as the system's performance characteristics, e.g. in economics or medicine. GDSi is a backpropagation-based training scheme that produces networks transformable into an equivalent and concise set of IF-THEN rules. This is achieved by imposing penalty terms on the network parameters that adapt the network to the expressive power of this class of rules.


Robust Parameter Estimation and Model Selection for Neural Network Regression

Neural Information Processing Systems

In this paper, it is shown that the conventional back-propagation (BPP) algorithm for neural network regression is robust to lever(cid:173) ages (data with:n corrupted), but not to outliers (data with y corrupted). A robust model is to model the error as a mixture of normal distribution. The influence function for this mixture model is calculated and the condition for the model to be robust to outliers is given. EM algorithm [5] is used to estimate the parameter. The usefulness of model selection criteria is also discussed.


The "Softmax" Nonlinearity: Derivation Using Statistical Mechanics and Useful Properties as a Multiterminal Analog Circuit Element

Neural Information Processing Systems

We use mean-field theory methods from Statistical Mechanics to derive the "softmax" nonlinearity from the discontinuous winner(cid:173) take-all (WTA) mapping. We give two simple ways of implementing "soft max" as a multiterminal network element. One of these has a number of important network-theoretic properties. It is a recipro(cid:173) cal, passive, incrementally passive, nonlinear, resistive multitermi(cid:173) nal element with a content function having the form of information(cid:173) theoretic entropy. These properties should enable one to use this element in nonlinear RC networks with such other reciprocal el(cid:173) ements as resistive fuses and constraint boxes to implement very high speed analog optimization algorithms using a minimum of hardware.


Grammatical Inference by Attentional Control of Synchronization in an Oscillating Elman Network

Neural Information Processing Systems

We show how an "Elman" network architecture, constructed from recurrently connected oscillatory associative memory network mod(cid:173) ules, can employ selective "attentional" control of synchronization to direct the flow of communication and computation within the architecture to solve a grammatical inference problem. Previously we have shown how the discrete time "Elman" network algorithm can be implemented in a network completely described by continuous ordinary differential equations. The time steps (ma(cid:173) chine cycles) of the system are implemented by rhythmic variation (clocking) of a bifurcation parameter. In this architecture, oscilla(cid:173) tion amplitude codes the information content or activity of a mod(cid:173) ule (unit), whereas phase and frequency are used to "softwire" the network. Only synchronized modules communicate by exchang(cid:173) ing amplitude information; the activity of non-resonating modules contributes incoherent crosstalk noise.


Central and Pairwise Data Clustering by Competitive Neural Networks

Neural Information Processing Systems

Data clustering amounts to a combinatorial optimization problem to re(cid:173) duce the complexity of a data representation and to increase its precision. Central and pairwise data clustering are studied in the maximum en(cid:173) tropy framework. For central clustering we derive a set of reestimation equations and a minimization procedure which yields an optimal num(cid:173) ber of clusters, their centers and their cluster probabilities. A meanfield approximation for pairwise clustering is used to estimate assignment probabilities. A se1fconsistent solution to multidimensional scaling and pairwise clustering is derived which yields an optimal embedding and clustering of data points in a d-dimensional Euclidian space.