Statistical Learning
Negative Margin Matters: Understanding Margin in Few-shot Classification
Liu, Bin, Cao, Yue, Lin, Yutong, Li, Qi, Zhang, Zheng, Long, Mingsheng, Hu, Han
This paper introduces a negative margin loss to metric learning based few-shot learning methods. The negative margin loss significantly outperforms regular softmax loss, and achieves state-of-the-art accuracy on three standard few-shot classification benchmarks with few bells and whistles. These results are contrary to the common practice in the metric learning field, that the margin is zero or positive. To understand why the negative margin loss performs well for the few-shot classification, we analyze the discriminability of learned features w.r.t different margins for training and novel classes, both empirically and theoretically. We find that although negative margin reduces the feature discriminability for training classes, it may also avoid falsely mapping samples of the same novel class to multiple peaks or clusters, and thus benefit the discrimination of novel classes.
Obliviousness Makes Poisoning Adversaries Weaker
Garg, Sanjam, Jha, Somesh, Mahloujifar, Saeed, Mahmoody, Mohammad, Thakurta, Abhradeep
Poisoning attacks have emerged as a significant security threat to machine learning (ML) algorithms. It has been demonstrated that adversaries who make small changes to the training set, such as adding specially crafted data points, can hurt the performance of the output model. Most of these attacks require the full knowledge of training data or the underlying data distribution. In this paper we study the power of oblivious adversaries who do not have any information about the training set. We show a separation between oblivious and full-information poisoning adversaries. Specifically, we construct a sparse linear regression problem for which LASSO estimator is robust against oblivious adversaries whose goal is to add a non-relevant features to the model with certain poisoning budget. On the other hand, non-oblivious adversaries, with the same budget, can craft poisoning examples based on the rest of the training data and successfully add non-relevant features to the model.
A lower bound for the ELBO of the Bernoulli Variational Autoencoder
Sicks, Robert, Korn, Ralf, Schwaar, Stefanie
We consider a variational autoencoder (VAE) for binary data. Our main innovations are an interpretable lower bound for its training objective, a modified initialization and architecture of such a VAE that leads to faster training, and a decision support for finding the appropriate dimension of the latent space via using a PCA. Numerical examples illustrate our theoretical result and the performance of the new architecture.
BayesFlow: Learning complex stochastic models with invertible neural networks
Radev, Stefan T., Mertens, Ulf K., Voss, Andreass, Ardizzone, Lynton, Kรถthe, Ullrich
Estimating the parameters of mathematical models is a common problem in almost all branches of science. However, this problem can prove notably difficult when processes and model descriptions become increasingly complex and an explicit likelihood function is not available. With this work, we propose a novel method for globally amortized Bayesian inference based on invertible neural networks which we call BayesFlow. The method uses simulation to learn a global estimator for the probabilistic mapping from observed data to underlying model parameters. A neural network pre-trained in this way can then, without additional training or optimization, infer full posteriors on arbitrary many real data sets involving the same model family. In addition, our method incorporates a summary network trained to embed the observed data into maximally informative summary statistics. Learning summary statistics from data makes the method applicable to modeling scenarios where standard inference techniques with hand-crafted summary statistics fail. We demonstrate the utility of BayesFlow on challenging intractable models from population dynamics, epidemiology, cognitive science and ecology. We argue that BayesFlow provides a general framework for building reusable Bayesian parameter estimation machines for any process model from which data can be simulated.
AirRL: A Reinforcement Learning Approach to Urban Air Quality Inference
Zhong, Huiqiang, Yin, Cunxiang, Wu, Xiaohui, Luo, Jinchang, He, JiaWei
Urban air pollution has become a major environmental problem that threatens public health. It has become increasingly important to infer fine-grained urban air quality based on existing monitoring stations. One of the challenges is how to effectively select some relevant stations for air quality inference. In this paper, we propose a novel model based on reinforcement learning for urban air quality inference. The model consists of two modules: a station selector and an air quality regressor. The station selector dynamically selects the most relevant monitoring stations when inferring air quality. The air quality regressor takes in the selected stations and makes air quality inference with deep neural network. We conduct experiments on a real-world air quality dataset and our approach achieves the highest performance compared with several popular solutions, and the experiments show significant effectiveness of proposed model in tackling problems of air quality inference.
A Survey on Edge Intelligence
Xu, Dianlei, Li, Tong, Li, Yong, Su, Xiang, Tarkoma, Sasu, Hui, Pan
Edge intelligence refers to a set of connected systems and devices for data collection, caching, processing, and analysis in locations close to where data is captured based on artificial intelligence. The aim of edge intelligence is to enhance the quality and speed of data processing and protect the privacy and security of the data. Although recently emerged, spanning the period from 2011 to now, this field of research has shown explosive growth over the past five years. In this paper, we present a thorough and comprehensive survey on the literature surrounding edge intelligence. We first identify four fundamental components of edge intelligence, namely edge caching, edge training, edge inference, and edge offloading, based on theoretical and practical results pertaining to proposed and deployed systems. We then aim for a systematic classification of the state of the solutions by examining research results and observations for each of the four components and present a taxonomy that includes practical problems, adopted techniques, and application goals. For each category, we elaborate, compare and analyse the literature from the perspectives of adopted techniques, objectives, performance, advantages and drawbacks, etc. This survey article provides a comprehensive introduction to edge intelligence and its application areas. In addition, we summarise the development of the emerging research field and the current state-of-the-art and discuss the important open issues and possible theoretical and technical solutions.
Why SMEs should embrace machine learning
RECENT technological advancements have placed artificial intelligence (AI) along with its subfield of machine learning (ML), at the forefront of transforming small and medium-sized enterprises (SMEs) digitally. Some SMEs are starting to tap into ML to shape their business processes and decision-making with the ultimate aim of raising profitability through revenue improvement, cost reduction and new sources of value creation. ML is seen as a continuation of the concepts around predictive analytics. However, a key difference in ML is that it uses mathematical algorithms to train computers in the processing and analysing of large amounts of data, allowing them to produce rules, identify patterns and generate classification predictions. It is important to note that computers automatically learn without human intervention or being explicitly programmed.
Machine Learning Using SAS Viya
Learn the theoretical foundation for different techniques associated with supervised machine learning models. You'll develop a series of supervised learning models including decision tree, ensemble of trees (forest and gradient boosting), neural networks and support vector machines. Demonstrations and exercises will reinforce all the concepts and the analytical approach to solving business problems. A business case study will guide you through all steps of the analytical life cycle, from problem understanding to model deployment, through data preparation, feature selection, model training and validation, and model assessment.
Evolutionary dataset optimisation: learning algorithm quality through evolution
This work presents a novel approach to learning the quality and performance of an algorithm through the use of evolution. When an algorithm is developed to solve a given problem, the designer is presented with questions about the performance of their proposed method and its relative performance against existing methods. This is an inherently difficult task. However, under the current paradigm, the standard response to this situation is to use a known fixed set of datasets - or simulate new datasets themselves - and a common metric amongst the proposed method and its competitors. The collated algorithms are then assessed based on this metric with often minimal consideration for the appropriateness or reliability of the datasets being used, and the robustness of the method(s) in question [1, 13, 19].
Put Your Money Where Your Strategy Is: Using Machine Learning to Analyze the Pentagon Budget - War on the Rocks
A "masterpiece" is how then-Deputy Defense Secretary Patrick Shanahan infamously described the Fiscal Year 2020 budget request. It would, he said, align defense spending with the U.S. National Defense Strategy -- both funding the future capabilities necessary to maintain an advantage over near-peer powers Russia and China, and maintaining readiness for ongoing counter-terror campaigns. While research and development funding increased in 2020, it did not represent the funding shift toward future capabilities that observers expected. Despite its massive size, the budget was insufficient to address the department's long-term challenges. Key emerging technologies identified by the department -- such as hypersonic weapons, artificial intelligence, quantum technologies, and directed-energy weapons -- still lacked a "clear and sustained commitment to investment." It was clear that the Department of Defense did not make the difficult tradeoffs necessary to fund long-term modernization.