Genre
Finding Alternate Features in Lasso
Hara, Satoshi, Maehara, Takanori
We propose a method for finding alternate features missing in the Lasso optimal solution. In ordinary Lasso problem, one global optimum is obtained and the resulting features are interpreted as task-relevant features. However, this can overlook possibly relevant features not selected by the Lasso. With the proposed method, we can provide not only the Lasso optimal solution but also possible alternate features to the Lasso solution. We show that such alternate features can be computed efficiently by avoiding redundant computations. We also demonstrate how the proposed method works in the 20 newsgroup data, which shows that reasonable features are found as alternate features.
Piecewise Deterministic Markov Processes for Continuous-Time Monte Carlo
Fearnhead, Paul, Bierkens, Joris, Pollock, Murray, Roberts, Gareth O
Monte Carlo methods, such as MCMC and SMC, have been central to the application of Bayesian statistics to real-world problems (Robert and Casella, 2011; McGrayne, 2011). These established Monte Carlo methods are based upon simulating discrete-time Markov processes. For example MCMC algorithms simulate a discrete-time Markov chain constructed to have a target distribution of interest, the posterior distribution in Bayesian inference, as its stationary distribution. Whilst SMC methods involve propagating and re-weighting particles so that a final set of weighted particles approximate a target distribution. The propagation step here also involves simulating from a discrete-time Markov chain. 1 In the past few years there have been exciting developments in MCMC and SMC methods based on continuoustime versions of these Monte Carlo methods. For example, continuous-time MCMC algorithms have been proposed (Peters and de With, 2012; Bouchard-Cรดtรฉ et al., 2015; Bierkens and Roberts, 2015; Bierkens et al., 2016) that involve simulating a continuous-time Markov process that has been designed to have a target distribution of interest as its stationary distribution. These continuous-time MCMC algorithms were originally motivated as they are examples of nonreversible Markov processes. There is substantial evidence that nonreversible MCMC algorithms will be more efficient than standard MCMC algorithms that are reversible (Neal, 1998; Diaconis et al., 2000; Neal, 2004; Bierkens, 2015), and there is empirical evidence that these continuous-time MCMC algorithms are more efficient than their discrete-time counterparts (see e.g.
Infinite Variational Autoencoder for Semi-Supervised Learning
Abbasnejad, Ehsan, Dick, Anthony, Hengel, Anton van den
This paper presents an infinite variational autoencoder (VAE) whose capacity adapts to suit the input data. This is achieved using a mixture model where the mixing coefficients are modeled by a Dirichlet process, allowing us to integrate over the coefficients when performing inference. Critically, this then allows us to automatically vary the number of autoencoders in the mixture based on the data. Experiments show the flexibility of our method, particularly for semi-supervised learning, where only a small number of training samples are available.
Tunable Sensitivity to Large Errors in Neural Network Training
Keren, Gil, Sabato, Sivan, Schuller, Bjรถrn
When humans learn a new concept, they might ignore examples that they cannot make sense of at first, and only later focus on such examples, when they are more useful for learning. We propose incorporating this idea of tunable sensitivity for hard examples in neural network learning, using a new generalization of the cross-entropy gradient step, which can be used in place of the gradient in any gradient-based training method. The generalized gradient is parameterized by a value that controls the sensitivity of the training process to harder training examples. We tested our method on several benchmark datasets. We propose, and corroborate in our experiments, that the optimal level of sensitivity to hard example is positively correlated with the depth of the network. Moreover, the test prediction error obtained by our method is generally lower than that of the vanilla cross-entropy gradient learner. We therefore conclude that tunable sensitivity can be helpful for neural network learning.
Learning Cost-Effective and Interpretable Regimes for Treatment Recommendation
Lakkaraju, Himabindu, Rudin, Cynthia
Decision makers, such as doctors, make crucial decisions su ch as recommending treatments to patients on a daily basis. Such decisions typically involve careful assessment of the subject's condition, analyzing the costs associated with the possible actions, and the nature of the consequent outcomes. Further, there might be costs associated with the assessmen t of the subject's condition itself (e.g., physical pain endured during medical tests, monetary costs etc.). Decision makers often leverage personal experience to make decisions in these contexts, wi thout considering data, even if massive amounts of it exist. Machine learning models could be of immense help in such scenarios - but these models would need to consider all three aspects discussed ab ove: predictions of counterfactuals, costs of gathering information, and costs of treatments. Fu rther, these models must be interpretable in order to create any reasonable chance of a human decision m aker actually using them. In this work, we address the problem of learning such cost-effectiv e, interpretable treatment regimes from observational data. Prior research addresses various aspects of the problem at h and in isolation. For instance, there exists a large body of literature on estimating treatment ef fects [5, 12, 4], recommending optimal treatments [1, 15, 6], and learning intelligible models for prediction [9, 7, 10, 2].
Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization
Li, Lisha, Jamieson, Kevin, DeSalvo, Giulia, Rostamizadeh, Afshin, Talwalkar, Ameet
Performance of machine learning algorithms depends critically on identifying a good set of hyperparameters. While current methods offer efficiencies by adaptively choosing new configurations to train, an alternative strategy is to adaptively allocate resources across the selected configurations. We formulate hyperparameter optimization as a pure-exploration non-stochastic infinitely many armed bandit problem where a predefined resource like iterations, data samples, or features is allocated to randomly sampled configurations. We introduce Hyperband for this framework and analyze its theoretical properties, providing several desirable guarantees. Furthermore, we compare Hyperband with state-of-the-art methods on a suite of hyperparameter optimization problems. We observe that Hyperband provides five times to thirty times speedup over state-of-the-art Bayesian optimization algorithms on a variety of deep-learning and kernel-based learning problems.
Parsimonious modeling with Information Filtering Networks
Barfuss, Wolfram, Massara, Guido Previde, Di Matteo, T., Aste, Tomaso
We introduce a methodology to construct parsimonious probabilistic models. This method makes use of Information Filtering Networks to produce a robust estimate of the global sparse inverse covariance from a simple sum of local inverse covariances computed on small sub-parts of the network. Being based on local and low-dimensional inversions, this method is computationally very efficient and statistically robust even for the estimation of inverse covariance of high-dimensional, noisy and short time-series. Applied to financial data our method results computationally more efficient than state-of-the-art methodologies such as Glasso producing, in a fraction of the computation time, models that can have equivalent or better performances but with a sparser inference structure. We also discuss performances with sparse factor models where we notice that relative performances decrease with the number of factors. The local nature of this approach allows us to perform computations in parallel and provides a tool for dynamical adaptation by partial updating when the properties of some variables change without the need of recomputing the whole model. This makes this approach particularly suitable to handle big datasets with large numbers of variables. Examples of practical application for forecasting, stress testing and risk allocation in financial systems are also provided.
On Design Mining: Coevolution and Surrogate Models
Preen, Richard J., Bull, Larry
Design mining [54, 55, 56] is the use of computational intelligence techniques to iteratively search and model the attribute space of physical objects evaluated directly through rapid prototyping to meet given objectives. It enables the exploitation of novel materials and processes without formal models or complex simulation, whilst harnessing the creativity of both computational and human design methods. A sample-model-search-sample loop creates an agile/flexible approach, i.e., primarily test-driven, enabling a continuing process of prototype design consideration and criteria refinement by both producers and users. Computational intelligence techniques have long been used in design, particularly for optimisation within simulations/models. Recent developments in additive-layer manufacturing (3D printing) means that it is now possible to work with over a hundred different materials, from ceramics to cells.
Nasa's Spirit Mars rover may have spotted signs of life on the red planet in 2007
It's been five years since NASA ended the Spirit rover's mission, but now, researchers say the robot may have discovered traces of life during its Mars investigation. A team of geoscientists has discovered that silica deposits from a region on the red planet dubbed'Home Plate' closely resemble those that form in Chilean hot springs at El Tatio. On Earth, these complex finger-like structures arise from a combination of biological and non-biological activity, suggesting a similar process may have taken place on Mars. The researchers compared opaline silica structures found at Home Plate (on left) with those from hot spring discharge channels at El Tatio (on right). The silica deposits on Mars were discovered after Spirit's right front wheel failed in 2007, forcing the robot to drag it across the ground like a plow near Home Plate, an eroded deposit of volcanic ash.
This survey drone takes safety seriously
New Zealand-based drone manufacturer Altus Intelligence wants to make sure its US$39,000 survey drones don't end up as rubble. Most of Altus' customers use its flagship drone, the Long Range Extreme Weather (LRX), for construction and engineering surveying/mapping and expect a rugged, dependable vehicle to get the job done. That's where the LRX's three separate fail-safe systems come in. The first is a triple auto pilot design, meaning that if anything goes wrong with one of the GPS streams, the other two will take over. The LRX is also armed with eight staggered propellers.