Goto

Collaborating Authors

 Deep Learning


InferCode: Self-Supervised Learning of Code Representations by Predicting Subtrees

arXiv.org Artificial Intelligence

Building deep learning models on source code has found many successful software engineering applications, such as code search, code comment generation, bug detection, code migration, and so on. Current learning techniques, however, have a major drawback that these models are mostly trained on datasets labeled for particular downstream tasks, and code representations may not be suitable for other tasks. While some techniques produce representations from unlabeled code, they are far from satisfactory when applied to downstream tasks. Although certain techniques generate representations from unlabeled code when applied to downstream tasks they are far from satisfactory. This paper proposes InferCode to overcome the limitation by adapting the self-supervised learning mechanism to build source code model. The key novelty lies in training code representations by predicting automatically identified subtrees from the context of the ASTs. Subtrees in ASTs are treated with InferCode as the labels for training code representations without any human labeling effort or the overhead of expensive graph construction, and the trained representations are no longer tied to any specific downstream tasks or code units. We trained an InferCode model instance using the Tree-based CNN as the encoder of a large set of Java code and applied it to downstream unsupervised tasks such as code clustering, code clone detection, cross-language code search or reused under a transfer learning scheme to continue training the model weights for supervised tasks such as code classification and method name prediction. Compared to previous code learning techniques applied to the same downstream tasks, such as Code2Vec, Code2Seq, ASTNN, higher performance results are achieved using our pre-trained InferCode model with a significant margin for most tasks including those involving different programming languages.


StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language Modeling

arXiv.org Artificial Intelligence

There are two major classes of natural language grammars -- the dependency grammar that models one-to-one correspondences between words and the constituency grammar that models the assembly of one or several corresponded words. While previous unsupervised parsing methods mostly focus on only inducing one class of grammars, we introduce a novel model, StructFormer, that can induce dependency and constituency structure at the same time. To achieve this, we propose a new parsing framework that can jointly generate a constituency tree and dependency graph. Then we integrate the induced dependency relations into the transformer, in a differentiable manner, through a novel dependencyconstrained self-attention mechanism. Experimental results show that our model can achieve strong results on unsupervised constituency parsing, unsupervised dependency parsing, and masked language modeling at the same time. Human languages have a rich latent structure. This structure is multifaceted, with the two major classes of grammar being dependency and constituency structures. There have been an exciting breath of recent work that are targeted at learning this structure in a data-driven unsupervised fashion.


Amazon SageMaker Automatic Model Tuning: Scalable Black-box Optimization

arXiv.org Machine Learning

Tuning complex machine learning systems is challenging. Machine learning models typically expose a set of hyperparameters, be it regularization, architecture, or optimization parameters, whose careful tuning is critical to achieve good performance. To democratize access to such systems, it is essential to automate this tuning process. This paper presents Amazon SageMaker Automatic Model Tuning (AMT), a fully managed system for black-box optimization at scale. AMT finds the best version of a machine learning model by repeatedly training it with different hyperparameter configurations. It leverages either random search or Bayesian optimization to choose the hyperparameter values resulting in the best-performing model, as measured by the metric chosen by the user. AMT can be used with built-in algorithms, custom algorithms, and Amazon SageMaker pre-built containers for machine learning frameworks. We discuss the core functionality, system architecture and our design principles. We also describe some more advanced features provided by AMT, such as automated early stopping and warm-starting, demonstrating their benefits in experiments.


Exploring Neural Networks Quantization via Layer-Wise Quantization Analysis

arXiv.org Machine Learning

Quantization is an essential step in the efficient deployment of deep learning models and as such is an increasingly popular research topic. An important practical aspect that is not addressed in the current literature is how to analyze and fix fail cases where the use of quantization results in excessive degradation. In this paper, we present a simple analytic framework that breaks down overall degradation to its per layer contributions. We analyze many common networks and observe that a layer's contribution is determined by both intrinsic (local) factors - the distribution of the layer's weights and activations - and extrinsic (global) factors having to do with the the interaction with the rest of the layers. Layer-wise analysis of existing quantization schemes reveals local fail-cases of existing techniques which are not reflected when inspecting their overall performance. As an example, we consider ResNext26 on which SoTA post-training quantization methods perform poorly. We show that almost all of the degradation stems from a single layer. The same analysis also allows for local fixes - applying a common weight clipping heuristic only to this layer reduces degradation to a minimum while applying the same heuristic globally results in high degradation. More generally, layer-wise analysis allows for a more nuanced examination of how quantization affects the network, enabling the design of better performing schemes.


Variational Beam Search for Online Learning with Distribution Shifts

arXiv.org Machine Learning

We consider the problem of online learning in the presence of sudden distribution shifts as frequently encountered in applications such as autonomous navigation. Distribution shifts require constant performance monitoring and re-training. They may also be hard to detect and can lead to a slow but steady degradation in model performance. To address this problem we propose a new Bayesian meta-algorithm that can both (i) make inferences about subtle distribution shifts based on minimal sequential observations and (ii) accordingly adapt a model in an online fashion. The approach uses beam search over multiple change point hypotheses to perform inference on a hierarchical sequential latent variable modeling framework. Our proposed approach is model-agnostic, applicable to both supervised and unsupervised learning, and yields significant improvements over state-of-the-art Bayesian online learning approaches.


Phase Retrieval with Holography and Untrained Priors: Tackling the Challenges of Low-Photon Nanoscale Imaging

arXiv.org Machine Learning

Phase retrieval is the inverse problem of recovering a signal from magnitude-only Fourier measurements, and underlies numerous imaging modalities, such as Coherent Diffraction Imaging (CDI). A variant of this setup, known as holography, includes a reference object that is placed adjacent to the specimen of interest before measurements are collected. The resulting inverse problem, known as holographic phase retrieval, is well-known to have improved problem conditioning relative to the original. This innovation, i.e. Holographic CDI, becomes crucial at the nanoscale, where imaging specimens such as viruses, proteins, and crystals require low-photon measurements. This data is highly corrupted by Poisson shot noise, and often lacks low-frequency content as well. In this work, we introduce a dataset-free deep learning framework for holographic phase retrieval adapted to these challenges. The key ingredients of our approach are the explicit and flexible incorporation of the physical forward model into an automatic differentiation procedure, the Poisson log-likelihood objective function, and an optional untrained deep image prior. We perform extensive evaluation under realistic conditions. Compared to competing classical methods, our method recovers signal from higher noise levels and is more resilient to suboptimal reference design, as well as to large missing regions of low frequencies in the observations. To the best of our knowledge, this is the first work to consider a dataset-free machine learning approach for holographic phase retrieval.


Predicting Generalization in Deep Learning via Local Measures of Distortion

arXiv.org Machine Learning

We study generalization in deep learning by appealing to complexity measures originally developed in approximation and information theory. While these concepts are challenged by the high-dimensional and data-defined nature of deep learning, we show that simple vector quantization approaches such as PCA, GMMs, and SVMs capture their spirit when applied layer-wise to deep extracted features giving rise to relatively inexpensive complexity measures that correlate well with generalization performance. We discuss our results in 2020 NeurIPS PGDL challenge.


DeepCube's Deep Learning Acceleration Platform Wins Seven Industry Awa

#artificialintelligence

DeepCube, the award-winning deep learning pioneer, today announced that its software-based deep learning acceleration platform has been recognized as a winner in several recent, prominent AI awards programs. These awards celebrate the top AI innovations and leaders across the globe. DeepCube's inclusion validates the immense potential of its patented deep learning acceleration platform that dramatically improves performance, latency and usability of deep learning on intelligent edge devices and in data centers. "Enterprises across industries are enticed by the potential for AI to unlock business impact and efficiencies; however, real-world, edge and data center deployments of deep learning remain out of reach, due to the immense size, processing power and memory requirements of these models," said Dr. Eli David, Co-Founder and Chief Technology Officer, DeepCube. "It's a difficult technical challenge, but it's one we're committed to solving at DeepCube. In 2020, we've made significant strides โ€“ both for our business and for the industry as a whole."


Semantic Image Segmentation with DeepLabv3-pytorch

#artificialintelligence

We will be using opencv to interface a webcam for reading in input from our screens and we'll use matplotlib's pyplot module to render the processed video feed to output. If you have multiple webcams you could create multiple such objects by passing the appropriate index; by default nowadays, most monitors have one inbuilt camera which could be indexed at 0th position. Subsequently, opencv reads images in a BGR format but while rendering we need to show it in RGB format; so we've written a tiny function that captures a frame in realtime and converts it from BGR format to RGB format above. With this, we're set with the input preprocessing steps. Let's look at how we'll set the stage for output now.


Chest X-Rays with Artificial Intelligence Catches More Lung Cancer

#artificialintelligence

Lung cancer detection and radiologist performance can get a boost from an artificial intelligence (AI) algorithm that pinpoints previously un-detected cancers on chest X-rays. In a study published in the Dec. 10 Radiology: Cardiothoracic Imaging, investigators from Seoul National University Hospital outlined how a commercially available deep-learning algorithm outperformed four thoracic radiologists on both first and second reads. Overall, said the team led by Ju Gang Nam, M.D., the algorithm offered both higher sensitivity and higher specificity, and it improved providers' performance as a seconder reader, leading to significantly improved detection rates. But, to date, the team said, adoption of computer-aided detection with chest X-ray has been slow because many providers still have lingering questions about whether it can perform well enough in clinical practice. To answer that question, Nam's team used an enriched dataset of 50 normal chest X-rays, as well as 168 posteroanterior chest X-rays with lung cancers.