Goto

Collaborating Authors

 Deep Learning


Finite Versus Infinite Neural Networks: an Empirical Study

arXiv.org Machine Learning

We perform a careful, thorough, and large scale empirical study of the correspondence between wide neural networks and kernel methods. By doing so, we resolve a variety of open questions related to the study of infinitely wide neural networks. Our experimental results include: kernel methods outperform fully-connected finite-width networks, but underperform convolutional finite width networks; neural network Gaussian process (NNGP) kernels frequently outperform neural tangent (NT) kernels; centered and ensembled finite networks have reduced posterior variance and behave more similarly to infinite networks; weight decay and the use of a large learning rate break the correspondence between finite and infinite networks; the NTK parameterization outperforms the standard parameterization for finite width networks; diagonal regularization of kernels acts similarly to early stopping; floating point precision limits kernel performance beyond a critical dataset size; regularized ZCA whitening improves accuracy; finite network performance depends non-monotonically on width in ways not captured by double descent phenomena; equivariance of CNNs is only beneficial for narrow networks far from the kernel regime. Our experiments additionally motivate an improved layer-wise scaling for weight decay which improves generalization in finite-width networks. Finally, we develop improved best practices for using NNGP and NT kernels for prediction, including a novel ensembling technique. Using these best practices we achieve state-of-the-art results on CIFAR-10 classification for kernels corresponding to each architecture class we consider.


A Hybrid Deep Learning Model for Predictive Flood Warning and Situation Awareness using Channel Network Sensors Data

arXiv.org Machine Learning

The objective of this study is to create and test a hybrid deep learning model, FastGRNN-FCN (Fast, Accurate, Stable and Tiny Gated Recurrent Neural Network-Fully Convolutional Network), for urban flood prediction and situation awareness using channel network sensors data. The study used Harris County, Texas as the testbed, and obtained channel sensor data from three historical flood events (e.g., 2016 Tax Day Flood, 2016 Memorial Day flood, and 2017 Hurricane Harvey Flood) for training and validating the hybrid deep learning model. The flood data are divided into a multivariate time series and used as the model input. Each input comprises nine variables, including information of the studied channel sensor and its predecessor and successor sensors in the channel network. Precision-recall curve and F-measure are used to identify the optimal set of model parameters. The optimal model with a weight of 1 and a critical threshold of 0.59 are obtained through one hundred iterations based on examining different weights and thresholds. The test accuracy and F-measure eventually reach 97.8% and 0.792, respectively. The model is then tested in predicting the 2019 Imelda flood in Houston and the results show an excellent match with the empirical flood. The results show that the model enables accurate prediction of the spatial-temporal flood propagation and recession and provides emergency response officials with a predictive flood warning tool for prioritizing the flood response and resource allocation strategies.


Hypergraph Learning with Line Expansion

arXiv.org Machine Learning

Previous hypergraph expansions are solely carried out on either vertex level or hyperedge level, thereby missing the symmetric nature of data co-occurrence, and resulting in information loss. To address the problem, this paper treats vertices and hyperedges equally and proposes a new hypergraph formulation named the \emph{line expansion (LE)} for hypergraphs learning. The new expansion bijectively induces a homogeneous structure from the hypergraph by treating vertex-hyperedge pairs as "line nodes". By reducing the hypergraph to a simple graph, the proposed \emph{line expansion} makes existing graph learning algorithms compatible with the higher-order structure and has been proven as a unifying framework for various hypergraph expansions. We evaluate the proposed line expansion on five hypergraph datasets, the results show that our method beats SOTA baselines by a significant margin.


WoodFisher: Efficient Second-Order Approximation for Neural Network Compression

arXiv.org Machine Learning

Second-order information, in the form of Hessian- or Inverse-Hessian-vector products, is a fundamental tool for solving optimization problems. Recently, there has been significant interest in utilizing this information in the context of deep neural networks; however, relatively little is known about the quality of existing approximations in this context. Our work examines this question, identifies issues with existing approaches, and proposes a method called WoodFisher to compute a faithful and efficient estimate of the inverse Hessian. Our main application is to neural network compression, where we build on the classic Optimal Brain Damage/Surgeon framework. We demonstrate that WoodFisher significantly outperforms popular state-of-the-art methods for one-shot pruning. Further, even when iterative, gradual pruning is considered, our method results in a gain in test accuracy over the state-of-the-art approaches, for pruning popular neural networks (like ResNet-50, MobileNetV1) trained on standard image classification datasets such as ImageNet ILSVRC. We examine how our method can be extended to take into account first-order information, as well as illustrate its ability to automatically set layer-wise pruning thresholds and perform compression in the limited-data regime.


Supervised Domain Adaptation: A Graph Embedding Perspective and a Rectified Experimental Protocol

arXiv.org Machine Learning

The performance of machine learning models tends to suffer when the distributions of the training and test data differ. Domain Adaptation is the process of closing the distribution gap between datasets. In this paper, we show that Domain Adaptation methods using pair-wise relationships between source and target domain data can be formulated as a Graph Embedding in which the domain labels are incorporated into the structure of the intrinsic and penalty graphs. We analyse the loss functions of existing state-of-the-art Supervised Domain Adaptation methods and demonstrate that they perform Graph Embedding. Moreover, we highlight some generalisation and reproducibility issues related to the experimental setup commonly used to demonstrate the few-shot learning capabilities of these methods. We propose a rectified evaluation setup for more accurately assessing and comparing Supervised Domain Adaptation methods, and report experiments on the standard benchmark datasets Office31 and MNIST-USPS.


Toward Robustness and Privacy in Federated Learning: Experimenting with Local and Central Differential Privacy

arXiv.org Artificial Intelligence

Federated Learning (FL) allows multiple participants to collaboratively train machine learning models by keeping their datasets local and exchanging model updates. Recent work has highlighted weaknesses related to robustness and privacy in FL, including backdoor, membership and property inference attacks. In this paper, we investigate whether and how Differential Privacy (DP) can be used to defend against attacks targeting both robustness and privacy in FL. To this end, we present a first-of-its-kind experimental evaluation of Local and Central Differential Privacy (LDP/CDP), assessing their feasibility and effectiveness. We show that both LDP and CDP do defend against backdoor attacks, with varying levels of protection and utility, and overall more effectively than non-DP defenses. They also mitigate white-box membership inference attacks, which our work is the first to show. Neither, however, defend against property inference attacks, prompting the need for further research in this space. Overall, our work also provides a re-usable measurement framework to quantify the trade-offs between robustness/privacy and utility in differentially private FL.


Road To Machine Learning Mastery: Interview With Kaggle GM Vladimir Iglovikov

#artificialintelligence

"I did not have lines in the resume that showed my ML expertise. I did not have a Data Science industry experience or relevant papers. For this week's ML practitioner's series, Analytics India Magazine got in touch with Vladimir Iglovikov, an ex-Spetsnaz, theoretical physicist and also a Kaggle GrandMaster. In this exclusive interview, he shares valuable information from his journey in the world of data science. After a brief stint in Russian special forces, Iglovikov enrolled for the Master's programme in theoretical Physics at the St.Petersburg State University whose distinguished alumni include President Vladimir Putin. In September 2010, Iglovikov moved to California to pursue a PhD in Physics from UC Davis and on completion of the degree, he moved to Silicon Valley in the summer of 2015. Currently, Iglovikov works as Sr. Software Engineer at Lyft, a ride-sharing company that operates in the United States and Canada. His work is centered around building robust machine learning models for autonomous vehicles at Lyft, Level5. Post PhD, Iglovikov had two options in hand. One was to pursue postdoc, and the other was to get into the industry as a software engineer. His career took a new turn when one of his friends introduced him to the world of data science. "I attended a lecture where the presenter talked about Data Science as the 4th paradigm of scientific discovery.


Deep Learning Components from Scratch in Python

#artificialintelligence

A subreddit dedicated for learning machine learning. Feel free to share any educational resources of machine learning. Also, we are a beginner-friendly sub-reddit, so don't be afraid to ask questions! This can include questions that are non-technical, but still highly relevant to learning machine learning such as a systematic approach to a machine learning problem.


Drug Discovery with Graph Neural Networks -- part 3

#artificialintelligence

Explanations Techniques help us understand the model's behaviour. For example, explanation methods are used to visualize certain parts of the image or to see how it reacts to a certain input. It is a well-established field of machine learning that has many different techniques which can be applied to deep learning (e.g. However, there have been only a few attempts to create explanation methods for graph neural networks (GNNs). Most of the "reuse" methods that were developed in deep learning and try to apply them in the graph domain. If you would like to learn more about state-of-the-art research on explainable GNNs, I would highly recommend looking over my previous article.


Hikvision 8MP Acusense Review

#artificialintelligence

Hikvision's Acusense is more a solution than a camera and we're testing outside the SEN network using the DS-7732NI-I4-16 32 channel, with 16 PoE Ports, 256Mbps throughput, H.265, 4K, 1.5RU, featuring 4 x HDD Bays a 3TB HDD. This NVR features a 4-core processor and supports H.265 intelligent compression, which aims to reduce bandwidth and storage requirements by up to 50 per cent. Before we get into the specifications of the camera, it's worth pointing out that the key to this camera is Acusense technology โ€“ a deep learning algorithm able to distinguish pedestrians, vehicles and objects and report events based on rules around what they do. Video clips are sorted into categories โ€“ people and vehicles โ€“ users click one of these categories and use time or location data to quickly locate the clip they need, making searches faster, as filtration has already been applied to footage. Key to this solution is that once it's set up, the camera does this automatically, all the time, and it also filters out'noise' so if there's an event, you're not battling through a river of video.