Asia
Deep Neural Network Approximation Theory
Grohs, Philipp, Perekrestenko, Dmytro, Elbrächter, Dennis, Bölcskei, Helmut
Deep neural networks have become state-of-the-art technology for a wide range of practical machine learning tasks such as image classification, handwritten digit recognition, speech recognition, or game intelligence. This paper develops the fundamental limits of learning in deep neural networks by characterizing what is possible if no constraints on the learning algorithm and the amount of training data are imposed. Concretely, we consider information-theoretically optimal approximation through deep neural networks with the guiding theme being a relation between the complexity of the function (class) to be approximated and the complexity of the approximating network in terms of connectivity and memory requirements for storing the network topology and the associated quantized weights. The theory we develop educes remarkable universality properties of deep networks. Specifically, deep networks are optimal approximants for vastly different function classes such as affine systems and Gabor systems. This universality is afforded by a concurrent invariance property of deep networks to time-shifts, scalings, and frequency-shifts. In addition, deep networks provide exponential approximation accuracy i.e., the approximation error decays exponentially in the number of non-zero weights in the network of vastly different functions such as the squaring operation, multiplication, polynomials, sinusoidal functions, general smooth functions, and even one-dimensional oscillatory textures and fractal functions such as the Weierstrass function, both of which do not have any known methods achieving exponential approximation accuracy. In summary, deep neural networks provide information-theoretically optimal approximation of a very wide range of functions and function classes used in mathematical signal processing.
Tree Tensor Networks for Generative Modeling
Cheng, Song, Wang, Lei, Xiang, Tao, Zhang, Pan
Matrix product states (MPS), a tensor network designed for one-dimensional quantum systems, has been recently proposed for generative modeling of natural data (such as images) in terms of `Born machine'. However, the exponential decay of correlation in MPS restricts its representation power heavily for modeling complex data such as natural images. In this work, we push forward the effort of applying tensor networks to machine learning by employing the Tree Tensor Network (TTN) which exhibits balanced performance in expressibility and efficient training and sampling. We design the tree tensor network to utilize the 2-dimensional prior of the natural images and develop sweeping learning and sampling algorithms which can be efficiently implemented utilizing Graphical Processing Units (GPU). We apply our model to random binary patterns and the binary MNIST datasets of handwritten digits. We show that TTN is superior to MPS for generative modeling in keeping correlation of pixels in natural images, as well as giving better log-likelihood scores in standard datasets of handwritten digits. We also compare its performance with state-of-the-art generative models such as the Variational AutoEncoders, Restricted Boltzmann machines, and PixelCNN. Finally, we discuss the future development of Tensor Network States in machine learning problems.
Data Masking with Privacy Guarantees
Pham, Anh T., Ghosh, Shalini, Yegneswaran, Vinod
We study the problem of data release with privacy, where data is made available with privacy guarantees while keeping the usability of the data as high as possible --- this is important in health-care and other domains with sensitive data. In particular, we propose a method of masking the private data with privacy guarantee while ensuring that a classifier trained on the masked data is similar to the classifier trained on the original data, to maintain usability. We analyze the theoretical risks of the proposed method and the traditional input perturbation method. Results show that the proposed method achieves lower risk compared to the input perturbation, especially when the number of training samples gets large. We illustrate the effectiveness of the proposed method of data masking for privacy-sensitive learning on $12$ benchmark datasets.
Dynamic Online Gradient Descent with Improved Query Complexity: A Theoretical Revisit
Zhao, Yawei, Zhu, En, Liu, Xinwang, Yin, Jianping
We provide a new theoretical analysis framework to investigate online gradient descent in the dynamic environment. Comparing with the previous work, the new framework recovers the state-of-the-art dynamic regret, but does not require extra gradient queries for every iteration. Specifically, when functions are $\alpha$ strongly convex and $\beta$ smooth, to achieve the state-of-the-art dynamic regret, the previous work requires $O(\kappa)$ with $\kappa = \frac{\beta}{\alpha}$ queries of gradients at every iteration. But, our framework shows that the query complexity can be improved to be $O(1)$, which does not depend on $\kappa$. The improvement is significant for ill-conditioned problems because that their objective function usually has a large $\kappa$.
Expanding the Reach of Federated Learning by Reducing Client Resource Requirements
Caldas, Sebastian, Konečny, Jakub, McMahan, H. Brendan, Talwalkar, Ameet
Communication on heterogeneous edge networks is a fundamental bottleneck in Federated Learning (FL), restricting both model capacity and user participation. To address this issue, we introduce two novel strategies to reduce communication costs: (1) the use of lossy compression on the global model sent server-to-client; and (2) Federated Dropout, which allows users to efficiently train locally on smaller subsets of the global model and also provides a reduction in both client-to-server communication and local computation. We empirically show that these strategies, combined with existing compression approaches for client-to-server communication, collectively provide up to a $14\times$ reduction in server-to-client communication, a $1.7\times$ reduction in local computation, and a $28\times$ reduction in upload communication, all without degrading the quality of the final model. We thus comprehensively reduce FL's impact on client device resources, allowing higher capacity models to be trained, and a more diverse set of users to be reached.
Comments on "Deep Neural Networks with Random Gaussian Weights: A Universal Classification Strategy?"
Gulcu, Talha Cihad, Gungor, Alper
In a recently published paper [1], it is shown that deep neural networks (DNNs) with random Gaussian weights preserve the metric structure of the data, with the property that the distance shrinks more when the angle between the two data points is smaller. We agree that the random projection setup considered in [1] preserves distances with a high probability. But as far as we are concerned, the relation between the angle of the data points and the output distances is quite the opposite, i.e., smaller angles result in a weaker distance shrinkage. This leads us to conclude that Theorem 3 and Figure 5 in [1] are not accurate. Hence the usage of random Gaussian weights in DNNs cannot provide an ability of universal classification or treating in-class and out-of-class data separately. Consequently, the behavior of networks consisting of random Gaussian weights only is not useful to explain how DNNs achieve state-of-art results in a large variety of problems.
Interpretable CNNs
Zhang, Quanshi, Wu, Ying Nian, Zhu, Song-Chun
This paper proposes a generic method to learn interpretable convolutional filters in a deep convolutional neural network (CNN), where each interpretable filter encodes features of a specific object part. Our method does not require additional annotations of object parts or textures for supervision. Instead, we use the same training data as traditional CNNs. Our method automatically assigns each interpretable filter in a high conv-layer with an object part of a certain category during the learning process. Such explicit knowledge representations in conv-layers of CNN help people clarify the logic encoded in the CNN, i.e., answering what patterns the CNN extracts from an input image and uses for prediction. We have tested our method using different benchmark CNNs with various structures to demonstrate the broad applicability of our method. Experiments have shown that our interpretable filters are much more semantically meaningful than traditional filters.
China's Huawei unveils chip for global big data market despite Western security warnings
BEIJING - Huawei Technologies Ltd. showed off a new processor chip for data centers and cloud computing Monday, expanding into new and growing markets despite Western warnings the company might be a security risk. Huawei and other Chinese technology companies that rely on Western technology are stepping up efforts to develop their own. The company based in southern China's Shenzhen has pushed ahead with commercial initiatives despite the Dec. 1 arrest of its chief financial officer, Meng Wanzhou, the daughter of Huawei founder Ren Zhengfei, in Canada on U.S. charges related to possible violations of trade sanctions on Iran. Huawei said the Kunpeng 920 chip is designed for servers that handle a flood of data from smartphones, video and other network services -- a fast growing sector with the development of artificial intelligence and the "internet of things." The company said it is part of a planned product lineup to support "intelligent computing."
Top 10 tech investments enterprise pros will make in 2019
The rapid speed of technology innovation means enterprise digital transformation efforts must keep pace with emerging technology trends, which can act as both disruptive threats and competitive opportunities, according to Altimeter's annual State of Digital Transformation report, released Thursday. The report surveyed 554 professionals from enterprise organizations across North America, Europe, and China. Digital transformation budgets rose dramatically over the past year, Altimeter found: Some 18% of companies reported digital transformation budgets ranging from $1 million to $5 million, while 17% said their budgets are between $5 million and $15 million. SEE: IT leader's guide to achieving digital transformation (Tech Pro Research) While 16% of companies reported budgets between $15 million and $30 million, this figure represents a 210% increase over the year before, according to the report. Budgets between $30 million and $50 million increased 234%, representing 13% of the companies surveyed.
CES 2019: Bread robots, bendy phones, talking toilets and everything else from day 1
An autonomous robot that can make bread, a phone that can transform into a tablet, and a talking toilet are among the gadgets on show at the world's biggest technology show. The BreadBot, made by US firm the Wilkinson Baking Company, can mix, proof and bake bread on its own and can produce a loaf every six minutes once up to full speed. It appeared at the CES Unveiled preview show, which offered an early glimpse of the gadgets going on display when the conference opens on Tuesday 8 January. The FlexPai foldable phone, built by Chinese startup Royole, is capable of folding from 0 to 180 degrees in order to function as both a tablet and a smartphone. Royole is billing the device as "the world's first commercial foldable smartphone", with a developer model currently available for pre-order at £1,209.