Goto

Collaborating Authors

 Europe


Deep Representation for Patient Visits from Electronic Health Records

arXiv.org Machine Learning

We show how to learn low-dimensional representations (embeddings) of patient visits from the corresponding electronic health record (EHR) where International Classification of Diseases (ICD) diagnosis codes are removed. We expect that these embeddings will be useful for the construction of predictive statistical models anticipated to drive personalized medicine and improve healthcare quality. These embeddings are learned using a deep neural network trained to predict ICD diagnosis categories. We show that our embeddings capture relevant clinical informations and can be used directly as input to standard machine learning algorithms like multi-output classifiers for ICD code prediction. We also show that important medical informations correspond to particular directions in our embedding space.


Scalable inference for crossed random effects models

arXiv.org Machine Learning

We analyze the complexity of Gibbs samplers for inference in crossed random effect models used in modern analysis of variance. We demonstrate that for certain designs the plain vanilla Gibbs sampler is not scalable, in the sense that its complexity is worse than proportional to the number of parameters and data. We thus propose a simple modification leading to a collapsed Gibbs sampler that is provably scalable. Although our theory requires some balancedness assumptions on the data designs, we demonstrate in simulated and real datasets that the rates it predicts match remarkably the correct rates in cases where the assumptions are violated. We also show that the collapsed Gibbs sampler, extended to sample further unknown hyperparameters, outperforms significantly alternative state of the art algorithms.


MOrdReD: Memory-based Ordinal Regression Deep Neural Networks for Time Series Forecasting

arXiv.org Machine Learning

Time series forecasting is ubiquitous in the modern world. Applications range from health care to astronomy, include climate modelling, financial trading and monitoring of critical engineering equipment. To offer value over this range of activities we must have models that not only provide accurate forecasts but that also quantify and adjust their uncertainty over time. Furthermore, such models must allow for multimodal, non-Gaussian behaviour that arises regularly in applied settings. In this work, we propose a novel, end-to-end deep learning method for time series forecasting. Crucially, our model allows the principled assessment of predictive uncertainty as well as providing rich information regarding multiple modes of future data values. Our approach not only provides an excellent predictive forecast, shadowing true future values, but also allows us to infer valuable information, such as the predictive distribution of the occurrence of critical events of interest, accurately and reliably even over long time horizons. We find the method outperforms other state-of-the-art algorithms, such as Gaussian Processes.


Revisiting First-Order Convex Optimization Over Linear Spaces

arXiv.org Machine Learning

Two popular examples of first-order optimization methods over linear spaces are coordinate descent and matching pursuit algorithms, with their randomized variants. While the former targets the optimization by moving along coordinates, the latter considers a generalized notion of directions. Exploiting the connection between the two algorithms, we present a unified analysis of both, providing affine invariant sublinear $\mathcal{O}(1/t)$ rates on smooth objectives and linear convergence on strongly convex objectives. As a byproduct of our affine invariant analysis of matching pursuit, our rates for steepest coordinate descent are the tightest known. Furthermore, we show the first accelerated convergence rate $\mathcal{O}(1/t^2)$ for matching pursuit on convex objectives.


Fr\'echet ChemblNet Distance: A metric for generative models for molecules

arXiv.org Machine Learning

The new wave of successful generative models in machine learning has increased the interest in deep learning driven de novo drug design. However, assessing the performance of such generative models is notoriously difficult. Metrics that are typically used to assess the performance of such generative models are the percentage of chemically valid molecules or the similarity to real molecules in terms of particular descriptors, such as the partition coefficient (logP) or druglikeness. However, method comparison is difficult because of the inconsistent use of evaluation metrics, the necessity for multiple metrics, and the fact that some of these measures can easily be tricked by simple rule-based systems. We propose a novel distance measure between two sets of molecules, called Fr\'echet ChemblNet distance (FCD), that can be used as an evaluation metric for generative models. The FCD is similar to a recently established performance metric for comparing image generation methods, the Fr\'echet Inception Distance (FID). Whereas the FID uses one of the hidden layers of InceptionNet, the FCD utilizes the penultimate layer of a deep neural network called "ChemblNet", which was trained to predict drug activities. Thus, the FCD metric takes into account chemically and biologically relevant information about molecules, and also measures the diversity of the set via the distribution of generated molecules. The FCD's advantage over previous metrics is that it can detect if generated molecules are a) diverse and have similar b) chemical and c) biological properties as real molecules. We further provide an easy-to-use implementation that only requires the SMILES representation of the generated molecules as input to calculate the FCD. Implementations are available at: https://www.github.com/bioinf-jku/FCD.


Counterfactual time-series prediction with encoder-decoder networks

arXiv.org Machine Learning

An important problem in the social sciences is estimating the effect of a policy intervention on an outcome over time. When interventions take place at an aggregate level (e.g., city or state), researchers make causal inferences by comparing the post-intervention outcomes for affected units ("treated") against the outcomes of a group of unaffected units ("control"). The synthetic control method (SCM) (Abadie, Diamond, and Hainmueller 2010) has become a popular method for making causal inferences on observational time-series. The method compares a single treated unit outcome with a synthetic control that combines the outcomes of multiple control units on the basis of their pre-intervention similarity with the treated unit. The SCM has several limitations.


Is AI smart enough to accurately measure an employee's value? – DXC Blogs

#artificialintelligence

Artificial intelligence (AI) and machine learning are supposed to help enterprises become more efficient by streamlining processes and taking over tasks currently performed by those unfocused, under-motivated humans known as employees. For those enterprise workers who don't lose their jobs to automation and machine learning, smart technologies can be used collaboratively by employees to make themselves more productive and effective. At least that's the theory, as reflected in comments by a couple of prominent CEOs appearing in a "future of work" panel at the recent World Economic Forum (WEF) in Davos, Switzerland. Bill McDermott, chief executive of Germany-based software giant SAP: "We should be optimistic. We see augmented humanity as the opportunity. It's there to enrich our lives, not take anything away from us."


Zuckerberg takes out ads to apologize as Facebook data misuse crisis intensifies

USATODAY - Tech Top Stories

A copy of'The Observer' shows an advertisement paid by Facebook in London, March 25, 2018. Facebook Chief Executive Mark Zuckerberg apologized for a "breach of trust" involving misused data from millions of Facebook users. The ads also appeared in The New York Times, Washington Post and Wall Street Journal. SAN FRANCISCO -- As Facebook continues to buffet winds of criticism, its founder took out full page ads in U.S. and British newspapers Sunday to apologize to consumers for not properly securing their personal data. "This was a breach of trust, and I'm sorry we didn't do more at the time," Mark Zuckerberg said in the signed ad, which was published in The New York Times, The Wall Street Journal, The Washington Post and six British papers.


Thanks to Translation Tech, Talking to Strangers Will Be Even Easier

WIRED

You totally prepared for this trip. You booked the flights months ago. You have all the best sights saved in Google Maps. But then you land in Munich or Kigali or Buenos Aires and realize you can't even identify the sign pointing toward baggage claim, much less tell your cabbie where you're headed. Luckily your phone can now do those things for you.


Financial Services Must Get More Out of Their Data - AI Holds The Key, says PwC

#artificialintelligence

With estimates showing that algorithmic trading systems handle 75% of the volume of global trades worldwide, financial services are feeling the impact of AI investment more than most. There's a number of reasons for this, not least because finance work is incredibly data-heavy. "Yesterday, I had an interesting conversation," offers Dr. Christian Westermann, Data & Analytics Partner with PwC Switzerland. He is better-placed than many to explain why AI is already such an area of interest within financial services. Formerly a space scientist, Westermann oversaw the early development of data analytics as a topic of interest at the firm from 2009 onwards.