Goto

Collaborating Authors

 Country


Linear interpolation gives better gradients than Gaussian smoothing in derivative-free optimization

arXiv.org Machine Learning

In this paper, we consider derivative free optimization problems, where the objective function is smooth but is computed with some amount of noise, the function evaluations are expensive and no derivative information is available. We are motivated by policy optimization problems in reinforcement learning that have recently become popular [Choromaski et al. 2018; Fazel et al. 2018; Salimans et al. 2016], and that can be formulated as derivative free optimization problems with the aforementioned characteristics. In each of these works some approximation of the gradient is constructed and a (stochastic) gradient method is applied. In [Salimans et al. 2016] the gradient information is aggregated along Gaussian directions, while in [Choromaski et al. 2018] it is computed along orthogonal direction. We provide a convergence rate analysis for a first-order line search method, similar to the ones used in the literature, and derive the conditions on the gradient approximations that ensure this convergence. We then demonstrate via rigorous analysis of the variance and by numerical comparisons on reinforcement learning tasks that the Gaussian sampling method used in [Salimans et al. 2016] is significantly inferior to the orthogonal sampling used in [Choromaski et al. 2018] as well as more general interpolation methods.


Generative Parameter Sampler For Scalable Uncertainty Quantification

arXiv.org Machine Learning

Uncertainty quantification has been a core of the statistical machine learning, but its computational bottleneck has been a serious challenge for both Bayesians and frequentists. We propose a model-based framework in quantifying uncertainty, called predictive-matching Generative Parameter Sampler (GPS). This procedure considers an Uncertainty Quantification (UQ) distribution, on the targeted parameter, which matches the corresponding predictive distribution to the observed data. This framework adopts a hierarchical modeling perspective such that each observation is modeled by an individual parameter. This individual parameterization permits the resulting inference to be computationally scalable and robust to outliers. Our approach is illustrated for linear models, Poisson processes, and deep neural networks for classification. The results show that the GPS is successful in providing uncertainty quantification as well as additional flexibility beyond what is allowed by classical statistical procedures under the postulated statistical models.


On approximating dropout noise injection

arXiv.org Machine Learning

This paper examines the assumptions of the derived equivalence between dropout noise injection and $L_2$ regularisation for logistic regression with negative log loss. We show that the approximation method is based on a divergent Taylor expansion, making, subsequent work using this approximation to compare the dropout trained logistic regression model with standard regularisers unfortunately ill-founded to date. Moreover, the approximation approach is shown to be invalid using any robust constraints. We show how this finding extends to general neural network topologies that use a cross-entropy prediction layer.


An Empirical Study on Hyperparameters and their Interdependence for RL Generalization

arXiv.org Artificial Intelligence

Recent results in Reinforcement Learning (RL) have shown that agents with limited training environments are susceptible to a large amount of overfitting across many domains. A key challenge for RL generalization is to quantitatively explain the effects of changing parameters on testing performance. Such parameters include architecture, regularization, and RL-dependent variables such as discount factor and action stochasticity. We provide empirical results that show complex and interdependent relationships between hyperparameters and generalization. We further show that several empirical metrics such as gradient cosine similarity and trajectory-dependent metrics serve to provide intuition towards these results.


Does It Make Sense? And Why? A Pilot Study for Sense Making and Explanation

arXiv.org Artificial Intelligence

Introducing common sense to natural language understanding systems has received increasing research attention. It remains a fundamental question on how to evaluate whether a system has a sense making capability. Existing benchmarks measures commonsense knowledge indirectly and without explanation. In this paper, we release a benchmark to directly test whether a system can differentiate natural language statements that make sense from those that do not make sense. In addition, a system is asked to identify the most crucial reason why a statement does not make sense. We evaluate models trained over large-scale language modeling tasks as well as human performance, showing that there are different challenges for system sense making.


Pre-training of Graph Augmented Transformers for Medication Recommendation

arXiv.org Artificial Intelligence

Medication recommendation is an important healthcare application. It is commonly formulated as a temporal prediction task. Hence, most existing works only utilize longitudinal electronic health records (EHRs) from a small number of patients with multiple visits ignoring a large number of patients with a single visit (selection bias). Moreover, important hierarchical knowledge such as diagnosis hierarchy is not leveraged in the representation learning process. To address these challenges, we propose G-BERT, a new model to combine the power of Graph Neural Networks (GNNs) and BERT (Bidirectional Encoder Representations from Transformers) for medical code representation and medication recommendation. We use GNNs to represent the internal hierarchical structures of medical codes. Then we integrate the GNN representation into a transformer-based visit encoder and pre-train it on EHR data from patients only with a single visit. The pre-trained visit encoder and representation are then fine-tuned for downstream predictive tasks on longitudinal EHRs from patients with multiple visits. G-BERT is the first to bring the language model pre-training schema into the healthcare domain and it achieved state-of-the-art performance on the medication recommendation task.


This Creepy AI Predicts What You Look Like Based on Your Voice

#artificialintelligence

A new artificial intelligence created by researchers at the Massachusetts Institute of Technology pulls off a staggering feat: by analyzing only a short audio clip of a person's voice, it reconstructs what they might look like in real life. The AI's results aren't perfect, but they're pretty good - a remarkable and somewhat terrifying example of how a sophisticated AI can make incredible inferences from tiny snippets of data. In a paper published this week to the preprint server arXiv, the team describes how it used trained a generative adversarial network to analyze short voice clips and "match several biometric characteristics of the speaker," resulting in "matching accuracies that are much better than chance." In practice, the Speech2Face algorithm seems to have an uncanny knack for spitting out rough likenesses of people based on nothing but their speaking voices. The MIT researchers urge caution on the project's GitHub page, acknowledging that the tech raises worrisome questions about privacy and discrimination.


DOM Pizza Checker

#artificialintelligence

Me! I am a world-first smart scanner that checks the quality of every Domino's pizza before it goes out the door. I've done my research and I know that nowadays it's ALL about looking good, so every pizza must be #nofilter Insta-worthy and meet our extremely high Quality Guarantee. With me in every store across Australia and New Zealand, product quality and consistency is about to go THROUGH THE ROOF! I sit above the cut bench – that's the area every pizza goes before being cut, boxed and delivered (not something in your gym class). I take a picture of the pizza and can recognise, analyse and grade pizzas based on pizza type, correct toppings and even whether the cheese is evenly spread!


Cracking open the black box of automated machine learning

#artificialintelligence

Researchers from MIT and elsewhere have developed an interactive tool that, for the first time, lets users see and control how automated machine-learning systems work. The aim is to build confidence in these systems and find ways to improve them. Designing a machine-learning model for a certain task -- such as image classification, disease diagnoses, and stock market prediction -- is an arduous, time-consuming process. Experts first choose from among many different algorithms to build the model around. Then, they manually tweak "hyperparameters" -- which determine the model's overall structure -- before the model starts training.


About 20 passengers injured as automated train in Yokohama travels in wrong direction, crashes into buffer

The Japan Times

YOKOHAMA - An automated train operated by Yokohama Seaside Line Co. on Saturday traveled in the wrong direction, causing about 20 people to be injured, a local fire department said. Some appeared to have suffered serious but non-life-threatening injuries as the train made contact with a buffer stop at Shin-Sugita Station, the department said, but other details were not immediately available. The trains are on an automated guideway transit system connecting Shin-Sugita and Kanazawa Hakkei in Yokohama.