Asia
Uncertainty Propagation in Deep Neural Network Using Active Subspace
Ji, Weiqi, Ren, Zhuyin, Law, Chung K.
The inputs of deep neural network (DNN) from real-world data usually come with uncertainties. Yet, it is challenging to propagate the uncertainty in the input features to the DNN predictions at a low computational cost. This work employs a gradient-based subspace method and response surface technique to accelerate the uncertainty propagation in DNN. Specifically, the active subspace method is employed to identify the most important subspace in the input features using the gradient of the DNN output to the inputs. Then the response surface within that low-dimensional subspace can be efficiently built, and the uncertainty of the prediction can be acquired by evaluating the computationally cheap response surface instead of the DNN models. In addition, the subspace can help explain the adversarial examples. The approach is demonstrated in MNIST datasets with a convolutional neural network.
Dynamic Demand Prediction for Expanding Electric Vehicle Sharing Systems: A Graph Sequence Learning Approach
Luo, Man, Wen, Hongkai, Luo, Yi, Du, Bowen, Klemmer, Konstantin, Zhu, Hongming
Electric Vehicle (EV) sharing systems have recently experienced unprecedented growth across the globe. During their fast expansion, one fundamental determinant for success is the capability of dynamically predicting the demand of stations as the entire system is evolving continuously. There are several challenges in this dynamic demand prediction problem. Firstly, unlike most of the existing work which predicts demand only for static systems or at few stages of expansion, in the real world we often need to predict the demand as or even before stations are being deployed or closed, to provide information and support for decision making. Secondly, for the stations to be deployed, there is no historical record or additional mobility data available to help the prediction of their demand. Finally, the impact of deploying/closing stations to the remaining stations in the system can be very complex. To address these challenges, in this paper we propose a novel dynamic demand prediction approach based on graph sequence learning, which is able to model the dynamics during the system expansion and predict demand accordingly. We use a local temporal encoding process to handle the available historical data at individual stations, and a dynamic spatial encoding process to take correlations between stations into account with graph convolutional neural networks. The encoded features are fed to a multi-scale prediction network, which forecasts both the long-term expected demand of the stations and their instant demand in the near future. We evaluate the proposed approach on real-world data collected from a major EV sharing platform in Shanghai for one year. Experimental results demonstrate that our approach significantly outperforms the state of the art, showing up to three-fold performance gain in predicting demand for the rapidly expanding EV sharing system.
meProp: Sparsified Back Propagation for Accelerated Deep Learning with Reduced Overfitting
Sun, Xu, Ren, Xuancheng, Ma, Shuming, Wang, Houfeng
We propose a simple yet effective technique for neural network learning. The forward propagation is computed as usual. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-$k$ elements (in terms of magnitude) are kept. As a result, only $k$ rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction ($k$ divided by the vector dimension) in the computational cost. Surprisingly, experimental results demonstrate that we can update only 1-4% of the weights at each back propagation pass. This does not result in a larger number of training iterations. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. The code is available at https://github.com/lancopku/meProp
Named Entity Recognition for Electronic Health Records: A Comparison of Rule-based and Machine Learning Approaches
Gorinski, Philip John, Wu, Honghan, Grover, Claire, Tobin, Richard, Talbot, Conn, Whalley, Heather, Sudlow, Cathie, Whiteley, William, Alex, Beatrice
This work investigates multiple approaches to Named Entity Recognition (NER) for text in Electronic Health Record (EHR) data. In particular, we look into the application of (i) rule-based, (ii) deep learning and (iii) transfer learning systems for the task of NER on brain imaging reports with a focus on records from patients with stroke. We explore the strengths and weaknesses of each approach, develop rules and train on a common dataset, and evaluate each system's performance on common test sets of Scottish radiology reports from two sources (brain imaging reports in ESS -- Edinburgh Stroke Study data collected by NHS Lothian as well as radiology reports created in NHS Tayside). Our comparison shows that a hand-crafted system is the most accurate way to automatically label EHR, but machine learning approaches can provide a feasible alternative where resources for a manual system are not readily available.
Exploring OpenStreetMap Availability for Driving Environment Understanding
Zheng, Yang, Izzat, Izzat H., Hansen, John H. L.
With the great achievement of artificial intelligence, vehicle technologies have advanced significantly from human centric driving towards fully automated driving. An intelligent vehicle should be able to understand the driver's perception of the environment as well as controlling behavior of the vehicle. Since high digital map information has been available to provide rich environmental context about static roads, buildings and traffic infrastructures, it would be worthwhile to explore map data capability for driving task understanding. Alternative to commercial used maps, the OpenStreetMap (OSM) data is a free open dataset, which makes it unique for the exploration research. This study is focused on two tasks that leverage OSM for driving environment understanding. First, driving scenario attributes are retrieved from OSM elements, which are combined with vehicle dynamic signals for the driving event recognition. Utilizing steering angle changes and based on a Bi-directional Recurrent Neural Network (Bi-RNN), a driving sequence is segmented and classified as lane-keeping, lane-change-left, lane-change-right, turn-left, and turn-right events. Second, for autonomous driving perception, OSM data can be used to render virtual street views, represented as prior knowledge to fuse with vision/laser systems for road semantic segmentation. Five different types of road masks are generated from OSM, images, and Lidar points, and fused to characterize the drivable space at the driver's perspective. An alternative data-driven approach is based on a Fully Convolutional Network (FCN), OSM availability for deep learning methods are discussed to reveal potential usage on compensating street view images and automatic road semantic annotation.
Multinomial Random Forests: Fill the Gap between Theoretical Consistency and Empirical Soundness
Li, Yiming, Bai, Jiawang, Tang, Qingtao, Jiang, Yong, Li, Chun, Xia, Shutao
Random forests (RF) are one of the most widely used ensemble learning methods in classification and regression tasks. Despite its impressive performance, its theoretical consistency, which would ensure that its result converges to the optimum as the sample size increases, has been left far behind. Several consistent random forest variants have been proposed, yet all with relatively poor performance compared to the original random forests. In this paper, a novel RF framework named multinomial random forests (MRF) is proposed. In the MRF, an impurity-based multinomial distribution is constructed as the basis for the selection of a splitting point. This ensures that a certain degree of randomness is achieved while the overall quality of the trees is not much different from the original random forests. We prove the consistency of the MRF and demonstrate with multiple datasets that it performs similarly as the original random forests and better than existent consistent random forest variants for both classification and regression tasks.
A Cold War Is Brewing Between the U.S. and China Over 5G and AI
Most recently, Chinese telecom equipment and consumer electronics giant Huawei filed suit against the U.S. government on Thursday, alleging that a law passed last August banning the company's hardware is unconstitutional because it unfairly targets Huawei. The arms race between the two superpowers over the tech talent, physical infrastructure and industrial muscle--not to mention the IP that comes with the technologies--that will unlock these advancements has been increasingly evident in diplomatic machinations, public investments of varying size and outward rhetoric from state leaders, as each country's government has moved the technologies to the top of their respective economic agendas. These struggles are about more than economic opportunity alone. The shift to 5G networks, which promise to eventually reach speeds up to 100 times faster than what's currently available, could expand the role of wireless communications systems into everything from power grids to traffic, while AI is already facilitating the mass collection and processing of data that will power critical tools like autonomous driving and facial recognition. In that future state, control over telecom equipment and AI-powered data centers could be like holding the keys to a society.
AI Weekly: Google's federated learning gets its day in the sun
A lot of news made headlines this week at the third annual TensorFlow Dev Summit. New versions of TensorFlow, including TensorFlow 2.0 with tf.keras as a central API and TensorFlow Lite 1.0 for mobile devices, were released, as was a $150 Coral board for edge TPU applications. Speed optimization for AI on mobile devices and a cleanup of TensorFlow's cluttered APIs is more than cosmetic -- these changes will shape how developers and businesses train AI systems. But the news that caught my eye was the release of TensorFlow for federated learning. TensorFlow Federated will provide distributed machine learning for developers to train models across many mobile devices without data ever leaving those devices.
I Quit My Job To Protest My Company's Work On Building Killer Robots
When I joined the artificial intelligence company Clarifai in early 2017, you could practically taste the promise in the air. My colleagues were brilliant, dedicated, and committed to making the world a better place. We founded Clarifai 4 Good where we helped students and charities, and we donated our software to researchers around the world whose projects had a socially beneficial goal. We were determined to be the one AI company that took our social responsibility seriously. I never could have predicted that two years later, I would have to quit this job on moral grounds.