Goto

Collaborating Authors

 Asia


Learning Discrete and Continuous Factors of Data via Alternating Disentanglement

arXiv.org Machine Learning

We address the problem of unsupervised disentanglement of discrete and continuous explanatory factors of data. We first show a simple procedure for minimizing the total correlation of the continuous latent variables without having to use a discriminator network or perform importance sampling, via cascading the information flow in the $\beta$-vae framework. Furthermore, we propose a method which avoids offloading the entire burden of jointly modeling the continuous and discrete factors to the variational encoder by employing a separate discrete inference procedure. This leads to an interesting alternating minimization problem which switches between finding the most likely discrete configuration given the continuous factors and updating the variational encoder based on the computed discrete factors. Experiments show that the proposed method clearly disentangles discrete factors and significantly outperforms current disentanglement methods based on the disentanglement score and inference network classification score. The source code is available at https://github.com/snu-mllab/DisentanglementICML19.


Detecting Adversarial Examples and Other Misclassifications in Neural Networks by Introspection

arXiv.org Machine Learning

Despite having excellent performances for a wide variety of tasks, modern neural networks are unable to provide a reliable confidence value allowing to detect misclassifications. This limitation is at the heart of what is known as an adversarial example, where the network provides a wrong prediction associated with a strong confidence to a slightly modified image. Moreover, this overconfidence issue has also been observed for regular errors and out-of-distribution data. We tackle this problem by what we call introspection, i.e. using the information provided by the logits of an already pretrained neural network. We show that by training a simple 3-layers neural network on top of the logit activations, we are able to detect misclassifications at a competitive level.


Multi-Task Kernel Null-Space for One-Class Classification

arXiv.org Machine Learning

The one-class kernel spectral regression (OC-KSR), the regression-based formulation of the kernel null-space approach has been found to be an effective Fisher criterion-based methodology for one-class classification (OCC), achieving state-of-the-art performance in one-class classification while providing relatively high robustness against data corruption. This work extends the OC-KSR methodology to a multi-task setting where multiple one-class problems share information for improved performance. By viewing the multi-task structure learning problem as one of compositional function learning, first, the OC-KSR method is extended to learn multiple tasks' structure \textit{linearly} by posing it as an instantiation of the separable kernel learning problem in a vector-valued reproducing kernel Hilbert space where an output kernel encodes tasks' structure while another kernel captures input similarities. Next, a non-linear structure learning mechanism is proposed which captures multiple tasks' relationships \textit{non-linearly} via an output kernel. The non-linear structure learning method is then extended to a sparse setting where different tasks compete in an output composition mechanism, leading to a sparse non-linear structure among multiple problems. Through extensive experiments on different data sets, the merits of the proposed multi-task kernel null-space techniques are verified against the baseline as well as other existing multi-task one-class learning techniques.


On the minimax optimality and superiority of deep neural network learning over sparse parameter spaces

arXiv.org Machine Learning

Deep learning has been applied to various tasks in the field of machine learning and has shown superiority to other common procedures such as kernel methods. To provide a better theoretical understanding of the reasons for its success, we discuss the performance of deep learning and other methods on a nonparametric regression problem with a Gaussian noise. Whereas existing theoretical studies of deep learning have been based mainly on mathematical theories of well-known function classes such as H\"{o}lder and Besov classes, we focus on function classes with discontinuity and sparsity, which are those naturally assumed in practice. To highlight the effectiveness of deep learning, we compare deep learning with a class of linear estimators representative of a class of shallow estimators. It is shown that the minimax risk of a linear estimator on the convex hull of a target function class does not differ from that of the original target function class. This results in the suboptimality of linear methods over a simple but non-convex function class, on which deep learning can attain nearly the minimax-optimal rate. In addition to this extreme case, we consider function classes with sparse wavelet coefficients. On these function classes, deep learning also attains the minimax rate up to log factors of the sample size, and linear methods are still suboptimal if the assumed sparsity is strong. We also point out that the parameter sharing of deep neural networks can remarkably reduce the complexity of the model in our setting.


A framework for the extraction of Deep Neural Networks by leveraging public data

arXiv.org Machine Learning

Machine learning models trained on confidential datasets are increasingly being deployed for profit. Machine Learning as a Service (MLaaS) has made such models easily accessible to end-users. Prior work has developed model extraction attacks, in which an adversary extracts an approximation of MLaaS models by making black-box queries to it. However, none of these works is able to satisfy all the three essential criteria for practical model extraction: (i) the ability to work on deep learning models, (ii) the non-requirement of domain knowledge and (iii) the ability to work with a limited query budget. We design a model extraction framework that makes use of active learning and large public datasets to satisfy them. We demonstrate that it is possible to use this framework to steal deep classifiers trained on a variety of datasets from image and text domains. By querying a model via black-box access for its top prediction, our framework improves performance on an average over a uniform noise baseline by 4.70 for image tasks and 2.11 for text tasks respectively, while using only 30% (30,000 samples) of the public dataset at its disposal.


A Coupled Operational Semantics for Goals and Commitments

Journal of Artificial Intelligence Research

Commitments capture how an agent relates to another agent, whereas goals describe states of the world that an agent is motivated to bring about. Commitments are elements of the social state of a set of agents whereas goals are elements of the private states of individual agents. It makes intuitive sense that goals and commitments are understood as being complementary to each other. More importantly, an agent's goals and commitments ought to be coherent, in the sense that an agent's goals would lead it to adopt or modify relevant commitments and an agent's commitments would lead it to adopt or modify relevant goals. However, despite the intuitive naturalness of the above connections, they have not been adequately studied in a formal framework. This article provides a combined operational semantics for goals and commitments by relating their respective life cycles as a basis for how these concepts (1) cohere for an individual agent and (2) engender cooperation among agents.


Ensemble Model Patching: A Parameter-Efficient Variational Bayesian Neural Network

arXiv.org Machine Learning

Two main obstacles preventing the widespread adoption of variational Bayesian neural networks are the high parameter overhead that makes them infeasible on large networks, and the difficulty of implementation, which can be thought of as "programming overhead." MC dropout [Gal and Ghahramani, 2016] is popular because it sidesteps these obstacles. Nevertheless, dropout is often harmful to model performance when used in networks with batch normalization layers [Li et al., 2018], which are an indispensable part of modern neural networks. We construct a general variational family for ensemble-based Bayesian neural networks that encompasses dropout as a special case. We further present two specific members of this family that work well with batch normalization layers, while retaining the benefits of low parameter and programming overhead, comparable to non-Bayesian training. Our proposed methods improve predictive accuracy and achieve almost perfect calibration on a ResNet-18 trained with ImageNet.


Facing the Rigid Demand of Artificial Intelligence Customer Service, Chengdu Xiaoduo AI Technology Co., Ltd., Based in Chengdu Tianfu Software Park, Completes Series B Financing

#artificialintelligence

The funds raised in this round are to be used for recruiting more high-caliber talents and developing new technologies. Xiaoduo AI has been engaged in the AI customer service application scenario for many years, which is a highlighted project in Chengdu's AI industry. From the perspective of the current enterprise volume and technical background, it has the potential to become a hidden champion in the field of AI customer service. Founded in 2014, Xiaoduo AI adheres to the vision of "becoming the AI expert of enterprises and creating better communication and service with AI", and has been committed to improving the efficiency of the customer service industry by leveraging the AI technology. At present, Xiaoduo AI serves more than 20,000 governmental and enterprise clients, including Meituan, Zhuanzhuan, YouShop, EMS, Robam Electric Appliance, XGIMI, TmallGenie, 1919.cn,


Automation and AI; What's the Real Difference? - ReadWrite

#artificialintelligence

Automation is good for business. It means delegating manual, mundane administrative tasks that suck up valuable hours to software or machines, freeing up time for human employees to focus on more complex, challenging, and creative work. Its benefits are twofoldโ€“better working conditions and employee engagement, as well as improving the bottom line by cutting costs. Think of it as the modern-day equivalent of the cotton gin. Before the invention of the cotton gin, people had to separate the cotton from their seeds by hand manually.


US warns about alleged spying threat from Chinese-made drones

FOX News

The US government is warning businesses about the risks of using Chinese-made aerial drones on claims they may pose a spying threat. On Monday, the Department of Homeland Security issued an industry alert over the alleged spying dangers, according to CNN. The alert doesn't name a specific company, but one of the biggest drone manufacturers in the world is DJI, which is based in Shenzhen, China. The department is worried the drone technologies can collect information and secretly send it back to their manufacturers in China. If this occurs, the Chinese government has the power to compel the manufacturer to hand over all the acquired data.