Country
A Probabilistic Machine Learning Approach to Detect Industrial Plant Faults
Fault detection in industrial plants is a hot research area as more and more sensor data are being collected throughout the industrial process. Automatic data-driven approaches are widely needed and seen as a promising area of investment. This paper proposes an effective machine learning algorithm to predict industrial plant faults based on classification methods such as penalized logistic regression, random forest and gradient boosted tree. A fault's start time and end time are predicted sequentially in two steps by formulating the original prediction problems as classification problems. The algorithms described in this paper won first place in the Prognostics and Health Management Society 2015 Data Challenge.
Best-of-K Bandits
Simchowitz, Max, Jamieson, Kevin, Recht, Benjamin
This paper studies the Best-of-K Bandit game: At each time the player chooses a subset S among all N-choose-K possible options and observes reward max(X(i) : i in S) where X is a random vector drawn from a joint distribution. The objective is to identify the subset that achieves the highest expected reward with high probability using as few queries as possible. We present distribution-dependent lower bounds based on a particular construction which force a learner to consider all N-choose-K subsets, and match naive extensions of known upper bounds in the bandit setting obtained by treating each subset as a separate arm. Nevertheless, we present evidence that exhaustive search may be avoided for certain, favorable distributions because the influence of high-order order correlations may be dominated by lower order statistics. Finally, we present an algorithm and analysis for independent arms, which mitigates the surprising non-trivial information occlusion that occurs due to only observing the max in the subset. This may inform strategies for more general dependent measures, and we complement these result with independent-arm lower bounds.
New Optimisation Methods for Machine Learning
A thesis submitted for the degree of Doctor of Philosophy of The Australian National University. In this work we introduce several new optimisation methods for problems in machine learning. Our algorithms broadly fall into two categories: optimisation of finite sums and of graph structured objectives. The finite sum problem is simply the minimisation of objective functions that are naturally expressed as a summation over a large number of terms, where each term has a similar or identical weight. Such objectives most often appear in machine learning in the empirical risk minimisation framework in the non-online learning setting. The second category, that of graph structured objectives, consists of objectives that result from applying maximum likelihood to Markov random field models. Unlike the finite sum case, all the non-linearity is contained within a partition function term, which does not readily decompose into a summation. For the finite sum problem, we introduce the Finito and SAGA algorithms, as well as variants of each. For graph-structured problems, we take three complementary approaches. We look at learning the parameters for a fixed structure, learning the structure independently, and learning both simultaneously. Specifically, for the combined approach, we introduce a new method for encouraging graph structures with the "scale-free" property. For the structure learning problem, we establish SHORTCUT, a O(n^{2.5}) expected time approximate structure learning method for Gaussian graphical models. For problems where the structure is known but the parameters unknown, we introduce an approximate maximum likelihood learning algorithm that is capable of learning a useful subclass of Gaussian graphical models.
Classification and Reconstruction of High-Dimensional Signals from Low-Dimensional Features in the Presence of Side Information
Renna, Francesco, Wang, Liming, Yuan, Xin, Yang, Jianbo, Reeves, Galen, Calderbank, Robert, Carin, Lawrence, Rodrigues, Miguel R. D.
This paper offers a characterization of fundamental limits on the classification and reconstruction of high-dimensional signals from low-dimensional features, in the presence of side information. We consider a scenario where a decoder has access both to linear features of the signal of interest and to linear features of the side information signal; while the side information may be in a compressed form, the objective is recovery or classification of the primary signal, not the side information. The signal of interest and the side information are each assumed to have (distinct) latent discrete labels; conditioned on these two labels, the signal of interest and side information are drawn from a multivariate Gaussian distribution. With joint probabilities on the latent labels, the overall signal-(side information) representation is defined by a Gaussian mixture model. We then provide sharp sufficient and/or necessary conditions for these quantities to approach zero when the covariance matrices of the Gaussians are nearly low-rank. These conditions, which are reminiscent of the well-known Slepian-Wolf and Wyner-Ziv conditions, are a function of the number of linear features extracted from the signal of interest, the number of linear features extracted from the side information signal, and the geometry of these signals and their interplay. Moreover, on assuming that the signal of interest and the side information obey such an approximately low-rank model, we derive expansions of the reconstruction error as a function of the deviation from an exactly low-rank model; such expansions also allow identification of operational regimes where the impact of side information on signal reconstruction is most relevant. Our framework, which offers a principled mechanism to integrate side information in high-dimensional data problems, is also tested in the context of imaging applications.
Feature Selection with Annealing for Computer Vision and Big Data Learning
Barbu, Adrian, She, Yiyuan, Ding, Liangjing, Gramajo, Gary
Many computer vision and medical imaging problems are faced with learning from large-scale datasets, with millions of observations and features. In this paper we propose a novel efficient learning scheme that tightens a sparsity constraint by gradually removing variables based on a criterion and a schedule. The attractive fact that the problem size keeps dropping throughout the iterations makes it particularly suitable for big data learning. Our approach applies generically to the optimization of any differentiable loss function, and finds applications in regression, classification and ranking. The resultant algorithms build variable screening into estimation and are extremely simple to implement. We provide theoretical guarantees of convergence and selection consistency. In addition, one dimensional piecewise linear response functions are used to account for nonlinearity and a second order prior is imposed on these functions to avoid overfitting. Experiments on real and synthetic data show that the proposed method compares very well with other state of the art methods in regression, classification and ranking while being computationally very efficient and scalable.
Mapping Temporal Variables into the NeuCube for Improved Pattern Recognition, Predictive Modelling and Understanding of Stream Data
Tu, Enmei, Kasabov, Nikola, Yang, Jie
This paper proposes a new method for an optimized mapping of temporal variables, describing a temporal stream data, into the recently proposed NeuCube spiking neural network architecture. This optimized mapping extends the use of the NeuCube, which was initially designed for spatiotemporal brain data, to work on arbitrary stream data and to achieve a better accuracy of temporal pattern recognition, a better and earlier event prediction and a better understanding of complex temporal stream data through visualization of the NeuCube connectivity. The effect of the new mapping is demonstrated on three bench mark problems. The first one is early prediction of patient sleep stage event from temporal physiological data. The second one is pattern recognition of dynamic temporal patterns of traffic in the Bay Area of California and the last one is the Challenge 2012 contest data set. In all cases the use of the proposed mapping leads to an improved accuracy of pattern recognition and event prediction and a better understanding of the data when compared to traditional machine learning techniques or spiking neural network reservoirs with arbitrary mapping of the variables.
Less is More: Nystr\"om Computational Regularization
Rudi, Alessandro, Camoriano, Raffaello, Rosasco, Lorenzo
We study Nystr\"om type subsampling approaches to large scale kernel methods, and prove learning bounds in the statistical learning setting, where random sampling and high probability estimates are considered. In particular, we prove that these approaches can achieve optimal learning bounds, provided the subsampling level is suitably chosen. These results suggest a simple incremental variant of Nystr\"om Kernel Regularized Least Squares, where the subsampling level implements a form of computational regularization, in the sense that it controls at the same time regularization and computations. Extensive experimental analysis shows that the considered approach achieves state of the art performances on benchmark large scale datasets.
Machine Learning and Personal Genome Informatics Contribute to Happiness Sciences and Wellbeing Computing
Kido, Takashi (Riken Genesis Co., Ltd.) | Swan, Melanie (MS Futures Group)
Two big recent revolutions: machine learning technologies; such as โdeep learningโ in Artificial Intelligence (AI), and personal genome informatics in biomedical science, provide us with new opportunities for understanding human happiness. Our ongoing important challenges are to discover our own truly meaningful personal happiness with the aid of AI and personal genome technologies. We have been developing a personal genome information agent entitled MyFinder, which supports searching for our inherited talents and maximizes our potential for a meaningful life. In the MyFinder project, we have provided a crowd-sourced DIY (Do it yourself) genomics research platform and conducted various โcitizen scienceโ projects in health and wellness. In this paper, we discuss how machine learning technologies and personal genome informat-ics might contribute to happiness sciences. We introduce the โSocial Intelligence Genomics and Empathy-Building Studyโ and report the preliminary results of applying deep learning and six other machine learning algorithms for predicting social intelligence levels from nine SNPs genetic profiles. We dis-cuss the possibilities and limitations of applying machine learning technologies for personal happiness trait prediction. We also discuss future AI challenges in the context of wellbeing computing.
AI and the Mitigation of Error: A Thermodynamics of Teams
Lawless, William Frere (Paine College) | Sofge, Donald A. (Naval Research Laboratory)
Traditional theories of social models conceptualize teams as distributed processors, disregarding the interdependence necessary to multi-task. Yet, interdependence characterizes social behavior. Instead, traditional theory favor cooperation, a state of least entropy production (LEP), without understanding the causes, limits or consequences of cooperation. As a simple example of interdependence, foraging prey overgraze forests free of predators. In our model, interdependence creates uncertainty, tradeoffs and signals (e.g., prices, coordination, innovation). Unlike individuals, the ability of teams to multitask reflects a quantum-like entanglement that represents maximum entropy production (MEP) when solving the problems signaled by society to improve its welfare. Our model supports findings that evolution in nature is driven by the MEP from making intelligent choices. Exploiting interdependence improves team intelligence, improves performance and reduces the risk of human error; forced cooperation disorganizes it by increasing the risk of error; e.g., if team cooperation improves teamwork, widespread forced cooperation in an autocracy or bureaucracy reduces social intelligence by adding unnecessary noise to signals. In our model, competition between teams self-organizes outsiders willing to sort through the noise for signals of the choices that improve social welfare (e.g., teams in courtrooms; science; entertainment; sports; businesses). Social systems organized around competition (e.g., stronger signals from robust checks and balances) better control a society by more correctly sizing teams to solve problems with fewer errors compared to autocracies or bureaucracies. Overall, we predict, the density of MEP directed at solving problems in a society with the constraints imposed from strong checks and balances, yet able to freely self-organize its labor and capital within those constraints, is denser.
Incorporating Human Dimension in Autonomous Decision-Making on Moral and Ethical Issues
Indurkhya, Bipin (Jagiellonian University) | Misztal-Radecka, Joanna (Jagiellonian University)
As autonomous systems are becoming more and more pervasive, they often have to make decisions concerning moral and ethical values. There are many approaches to incorporating moral values in autonomous decision-making that are based on some sort of logical deduction. However, we argue here, in order for decision-making to seem persuasive to humans, it needs to reflect human values and judgments. Employing some insights from our ongoing researchusing features of the blackboard architecture for a context-aware recommender system, and a legal decision-making system that incorporates supra-legal aspects, we aim to explore if this architecture can also be adapted to implement a moral decision-making system that generates rationales that are persuasive to humans. Our vision is that such a system can be used as an advisory system to consider a situation from different moral perspectives, and generate ethical pros and cons of taking a particular course of action in a given context.