Goto

Collaborating Authors

 Deep Learning


Measuring Robustness to Natural Distribution Shifts in Image Classification

arXiv.org Machine Learning

We study how robust current ImageNet models are to distribution shifts arising from natural variations in datasets. Most research on robustness focuses on synthetic image perturbations (noise, simulated weather artifacts, adversarial examples, etc.), which leaves open how robustness on synthetic distribution shift relates to distribution shift arising in real data. Informed by an evaluation of 204 ImageNet models in 213 different test conditions, we find that there is often little to no transfer of robustness from current synthetic to natural distribution shift. Moreover, most current techniques provide no robustness to the natural distribution shifts in our testbed. The main exception is training on larger and more diverse datasets, which in multiple cases increases robustness, but is still far from closing the performance gaps. Our results indicate that distribution shifts arising in real data are currently an open research problem. We provide our testbed and data as a resource for future work at https://modestyachts.github.io/imagenet-testbed/ .


Relational Fusion Networks: Graph Convolutional Networks for Road Networks

arXiv.org Machine Learning

The application of machine learning techniques in the setting of road networks holds the potential to facilitate many important intelligent transportation applications. Graph Convolutional Networks (GCNs) are neural networks that are capable of leveraging the structure of a network. However, many implicit assumptions of GCNs do not apply to road networks. We introduce the Relational Fusion Network (RFN), a novel type of GCN designed specifically for road networks. In particular, we propose methods that outperform state-of-the-art GCNs by 21%-40% on two machine learning tasks in road networks. Furthermore, we show that state-of-the-art GCNs may fail to effectively leverage road network structure and may not generalize well to other road networks.


Sparsity Turns Adversarial: Energy and Latency Attacks on Deep Neural Networks

arXiv.org Machine Learning

Adversarial attacks have exposed serious vulnerabilities in Deep Neural Networks (DNNs) through their ability to force misclassifications through human-imperceptible perturbations to DNN inputs. We explore a new direction in the field of adversarial attacks by suggesting attacks that aim to degrade the computational efficiency of DNNs rather than their classification accuracy. Specifically, we propose and demonstrate sparsity attacks, which adversarial modify a DNN's inputs so as to reduce sparsity (or the presence of zero values) in its internal activation values. In resource-constrained systems, a wide range of hardware and software techniques have been proposed that exploit sparsity to improve DNN efficiency. The proposed attack increases the execution time and energy consumption of sparsity-optimized DNN implementations, raising concern over their deployment in latency and energy-critical applications. We propose a systematic methodology to generate adversarial inputs for sparsity attacks by formulating an objective function that quantifies the network's activation sparsity, and minimizing this function using iterative gradient-descent techniques. We launch both white-box and black-box versions of adversarial sparsity attacks on image recognition DNNs and demonstrate that they decrease activation sparsity by up to 1.82x. We also evaluate the impact of the attack on a sparsity-optimized DNN accelerator and demonstrate degradations up to 1.59x in latency, and also study the performance of the attack on a sparsity-optimized general-purpose processor. Finally, we evaluate defense techniques such as activation thresholding and input quantization and demonstrate that the proposed attack is able to withstand them, highlighting the need for further efforts in this new direction within the field of adversarial machine learning.


The Radicalization Risks of GPT-3 and Advanced Neural Language Models

arXiv.org Artificial Intelligence

In this paper, we expand on our previous research of the potential for abuse of generative language models by assessing GPT-3. Experimenting with prompts representative of different types of extremist narrative, structures of social interaction, and radical ideologies, we find that GPT-3 demonstrates significant improvement over its predecessor, GPT-2, in generating extremist texts. We also show GPT-3's strength in generating text that accurately emulates interactive, informational, and influential content that could be utilized for radicalizing individuals into violent far-right extremist ideologies and behaviors. While OpenAI's preventative measures are strong, the possibility of unregulated copycat technology represents significant risk for large-scale online radicalization and recruitment; thus, in the absence of safeguards, successful and efficient weaponization that requires little experimentation is likely. AI stakeholders, the policymaking community, and governments should begin investing as soon as possible in building social norms, public policy, and educational initiatives to preempt an influx of machine-generated disinformation and propaganda. Mitigation will require effective policy and partnerships across industry, government, and civil society.


On the Orthogonality of Knowledge Distillation with Other Techniques: From an Ensemble Perspective

arXiv.org Artificial Intelligence

To put a state-of-the-art neural network to practical use, it is necessary to design a model that has a good trade-off between the resource consumption and performance on the test set. Many researchers and engineers are developing methods that enable training or designing a model more efficiently. Developing an efficient model includes several strategies such as network architecture search, pruning, quantization, knowledge distillation, utilizing cheap convolution, regularization, and also includes any craft that leads to a better performance-resource trade-off. When combining these technologies together, it would be ideal if one source of performance improvement does not conflict with others. We call this property as the orthogonality in model efficiency. In this paper, we focus on knowledge distillation and demonstrate that knowledge distillation methods are orthogonal to other efficiency-enhancing methods both analytically and empirically. Analytically, we claim that knowledge distillation functions analogous to a ensemble method, bootstrap aggregating. This analytical explanation is provided from the perspective of implicit data augmentation property of knowledge distillation. Empirically, we verify knowledge distillation as a powerful apparatus for practical deployment of efficient neural network, and also introduce ways to integrate it with other methods effectively.


[Full text] Conceptualising Artificial Intelligence as a Digital Healthcare Innova

#artificialintelligence

Abstract: Artificial intelligence (AI) is widely recognised as a transformative innovation and is already proving capable of outperforming human clinicians in the diagnosis of specific medical conditions, especially in image analysis within dermatology and radiology. These abilities are enhanced by the capacity of AI systems to learn from patient records, genomic information and real-time patient data. Whilst AI research is mounting, less attention has been paid to the practical implications on healthcare services and potential barriers to implementation. AI is recognised as a "Software as a Medical Device (SaMD)" and is increasingly becoming a topic of interest for regulators. Unless the introduction of AI is carefully considered and gradual, there are risks of automation bias, overdependence and long-term staffing problems. This is in addition to already well-documented generic risks associated with AI, such as data privacy, algorithmic biases and corrigibility.


Neural network: what is a neural network?

#artificialintelligence

The amount of data needed to train deep neural networks can be truly immense, and in many commercial systems today it can take weeks or months of training to obtain adequate performance. This is even considering that networks are often trained in parallel across many highly optimized machines with specialized hardware such as ASICs and GPUs. This is a significant limitation of deep neural networks โ€“ one could imagine the frustration of a deep learning engineer who learns only after weeks of expensive training that there was an error or bug in the design of the neural network that inhibited performance.


Can GPT-3 Really Help You and Your Company?

#artificialintelligence

GPT-3 is a new text-generating program from OpenAI. This model is pre-trained, but it is never touched again. Specifically, they trained GPT-3 on a dataset of half a trillion words for 175 billion parameters, which is 10x more than any previous non-sparse language model. Then, there is no more fine-tuning to do with this model, only a few-shot demonstrations specified purely via text interaction with the model. The few-shot works by giving a certain amount of examples of context and completion, and then one final example of context, with the model expected to provide the completion without changing the model's parameters.


AI Weekly: What ML practitioners are doing about climate change

#artificialintelligence

A lot happened this week in the AI space. The Guardian wrote an article with GPT-3 and again demonstrated that no matter what OpenAI paid to train and create the language model, the free marketing might be worth more. After losing the JEDI cloud contract appeal with the Pentagon, Amazon appointed to its board Keith Alexander, who oversaw the National Security Agency mass surveillance revealed by Edward Snowden leaks in 2013. And Portland passed the strictest facial recognition bans in U.S. history, outlawing government and business use of the technology. However, AI Weekly attempts to reach into the zeitgeist and highlight the issues on people's minds. This week without question it's the smoke that has hung over the western United States and the underlying problem of climate change.


Comparison of Activation Functions for Deep Neural Networks

#artificialintelligence

Activation functions play a key role in neural networks, so it is essential to understand the advantages and disadvantages to achieve better performance. It is necessary to start by introducing the non-linear activation functions, which is an alternative to the best known sigmoid function. It is important to remember that many different conditions are important when evaluating the final performance of activation functions. It is necessary to draw attention to the importance of mathematics and the derivative process at this point. So, if you're ready, let's roll up the sleeves and get our hands dirty!