Goto

Collaborating Authors

 Asia


Grids versus Graphs: Partitioning Space for Improved Taxi Demand-Supply Forecasts

arXiv.org Machine Learning

Abstract--Accurate taxi demand-supply forecasting is a challenging applicationof ITS (Intelligent Transportation Systems), due to the complex spatial and temporal patterns. We investigate the impact of different spatial partitioning techniques on the prediction performance of an LSTM (Long Short-Term Memory) network, in the context of taxi demand-supply forecasting. We consider two tessellation schemes: (i) the variable-sized Voronoi tessellation, and (ii) the fixed-sized Geohash tessellation. While the widely employed ConvLSTM (Convolutional LSTM) can model fixed-sized Geohash partitions, the standard convolutional filters cannot be applied on the variable-sized Voronoi partitions. To explore the Voronoi tessellation scheme, we propose the use of GraphLSTM (Graph-based LSTM), by representing the Voronoi spatial partitions as nodes on an arbitrarily structured graph. The GraphLSTM offers competitive performance against ConvLSTM, atlower computational complexity, across three realworld large-scale taxi demand-supply data sets, with different performance metrics. To ensure superior performance across diverse settings, a HEDGE based ensemble learning algorithm is applied over the ConvLSTM and the GraphLSTM networks. I. INTRODUCTION Spatiotemporal forecasting has a wide range of applications, rangingfrom epidemic detection [1], energy management [2], to cellular traffic [3], among others. Location-based taxi demand and supply forecasting, one of the key components ofITS (Intelligent Transportation Systems), also relies heavily on accurate spatiotemporal forecasting. Mobility-on- Demand services such as e-hailing taxis, which have gained tremendous popularity in the recent years, often face taxi demand-supply imbalances. During peak and off-peak hours, mismatches occur between the spatial distributions of the taxi demand and the available drivers, resulting in either scarcity or abundance of vacant taxis. For example, Figure 1 presents a case of demand-supply mismatch averaged over all Mondays near the city center in Bengaluru, India.


Differentially Private Continual Learning

arXiv.org Machine Learning

Catastrophic forgetting can be a significant problem for institutions that must delete historic data for privacy reasons. For example, hospitals might not be able to retain patient data permanently. But neural networks trained on recent data alone will tend to forget lessons learned on old data. We present a differentially private continual learning framework based on variational inference. We estimate the likelihood of past data given the current model using differentially private generative models of old datasets.


A Unifying Bayesian View of Continual Learning

arXiv.org Machine Learning

Some machine learning applications require continual learning - where data comes in a sequence of datasets, each is used for training and then permanently discarded. From a Bayesian perspective, continual learning seems straightforward: Given the model posterior one would simply use this as the prior for the next task. However, exact posterior evaluation is intractable with many models, especially with Bayesian neural networks (BNNs). Instead, posterior approximations are often sought. Unfortunately, when posterior approximations are used, prior-focused approaches do not succeed in evaluations designed to capture properties of realistic continual learning use cases. As an alternative to prior-focused methods, we introduce a new approximate Bayesian derivation of the continual learning loss. Our loss does not rely on the posterior from earlier tasks, and instead adapts the model itself by changing the likelihood term. We call these approaches likelihood-focused. We then combine prior- and likelihood-focused methods into one objective, tying the two views together under a single unifying framework of approximate Bayesian continual learning.


Reinforcement Learning Without Backpropagation or a Clock

arXiv.org Machine Learning

Reinforcement learning (RL) algorithms share qualitative similarities with the algorithms implemented byanimal brains. However, there remain clear differences between these two types of algorithms. For example, while RL algorithms using artificial neural networks require information to flow backwards through the network via the backpropagation algorithm, there is currently debate about whether this is feasible in biological neural implementations (Werbos and Davis, 2016). Policy gradient coagent networks (PGCNs) are a class of RL algorithms that were introduced to remove this possibly biologically implausible property of RL algorithms--they use artificial neural networks but do not use the backpropagation algorithm (Thomas, 2011). Since their introduction, PGCN algorithms have proven to be not only a possible improvement in biological plausibility, but a practical tool for improving RL agents. They were used to solve RL problems with high-dimensional action spaces (Thomas and Barto, 2012), are the RL precursor to the more general stochastic computation graphs (Schulman et al., 2015), and, as we will show in this paper, generalize the recently proposed option-critic architecture (Bacon et al., 2017), while drastically simplifying key derivations.


Learning Simple Thresholded Features with Sparse Support Recovery

arXiv.org Machine Learning

The thresholded feature has recently emerged as an extremely efficient, yet rough empirical approximation, of the time-consuming sparse coding inference process. Such an approximation has not yet been rigorously examined, and standard dictionaries often lead to non-optimal performance when used for computing thresholded features. In this paper, we first present two theoretical recovery guarantees for the thresholded feature to exactly recover the nonzero support of the sparse code. Motivated by them, we then formulate the Dictionary Learning for Thresholded Features (DLTF) model, which learns an optimized dictionary for applying the thresholded feature. In particular, for the $(k, 2)$ norm involved, a novel proximal operator with log-linear time complexity $O(m\log m)$ is derived. We evaluate the performance of DLTF on a vast range of synthetic and real-data tasks, where DLTF demonstrates remarkable efficiency, effectiveness and robustness in all experiments. In addition, we briefly discuss the potential link between DLTF and deep learning building blocks.


Iterative Local Voting for Collective Decision-making in Continuous Spaces

Journal of Artificial Intelligence Research

Many societal decision problems lie in high-dimensional continuous spaces not amenable to the voting techniques common for their discrete or single-dimensional counterparts. These problems are typically discretized before running an election or decided upon through negotiation by representatives. We propose a algorithm called Iterative Local Voting for collective decision-making in this setting. In this algorithm, voters are sequentially sampled and asked to modify a candidate solution within some local neighborhood of its current value, as defined by a ball in some chosen norm, with the size of the ball shrinking at a specified rate. We first prove the convergence of this algorithm under appropriate choices of neighborhoods to Pareto optimal solutions with desirable fairness properties in certain natural settings: when the voters' utilities can be expressed in terms of some form of distance from their ideal solution, and when these utilities are additively decomposable across dimensions. In many of these cases, we obtain convergence to the societal welfare maximizing solution.We then describe an experiment in which we test our algorithm for the decision of the U.S. Federal Budget on Mechanical Turk with over 2,000 workers, employing neighborhoods defined by various L-Norm balls. We make several observations that inform future implementations of such a procedure.


Proper-Composite Loss Functions in Arbitrary Dimensions

arXiv.org Machine Learning

The study of a machine learning problem is in many ways is difficult to separate from the study of the loss function being used. One avenue of inquiry has been to look at these loss functions in terms of their properties as scoring rules via the proper-composite representation, in which predictions are mapped to probability distributions which are then scored via a scoring rule. However, recent research so far has primarily been concerned with analysing the (typically) finite-dimensional conditional risk problem on the output space, leaving aside the larger total risk minimisation. We generalise a number of these results to an infinite dimensional setting and in doing so we are able to exploit the familial resemblance of density and conditional density estimation to provide a simple characterisation of the canonical link.


South Korean tanker Stellar Daisy found on ocean floor 2 years after it sank, explorers say

FOX News

The Stellar Daisy, a massive South Korean tanker that sank in March 2017, was spotted on the floor of the South Atlantic Ocean nearly two years later, the CEO of an ocean exploration company revealed Sunday. This discovery could shed new light on exactly what caused the vessel to tilt and sink and provide some closure to the families of the 22 crew members who died. "We are pleased to report that we have located Stellar Daisy, in particular for our client, the South Korean Government, but also for the families of those who lost loved ones in this tragedy," Ocean Infinity CEO Oliver Plunkett said. "Through the deployment of multiple state of the art (autonomous underwater vehicles), we are covering the seabed with unprecedented speed and accuracy." The Stellar Daisy sank on March 31, 2017, nearly 2,500 miles east of Uruguay, while transporting iron ore from Brazil to China.


Google Software Engineer, Machine Learning Job in Tokyo

#artificialintelligence

Minimum qualifications: BA/BS degree in Computer Science or related technical field or equivalent practical experience. 2 years of work or educational experience in Machine Learning or Artificial Intelligence. 1 year of relevant work experience, including software development. Experience with one or more general purpose programming languages including but not limited to: Java, C/C or Python. Preferred qualifications: MS or PhD degree in Computer Science, Artificial Intelligence, Machine Learning, or related technical field. About the job Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search.


Basic AI/ML Safety Guide for the Non-Technical – Katie Evanko-Douglas – Medium

#artificialintelligence

Artificial Intelligence (AI) and Machine Learning (ML) affect and will continue to affect many aspects of our everyday lives at an increasing rate. Because it affects everybody, it is important for everybody to understand these issues at a basic level. The content of this blog post will likely seem a gross simplification to people who are deeply technical. It's purpose is to shed light on and make comprehensible at a high level the current and future issues we face as a species when it comes to Artificial Intelligence and Machine Learning. The Crash Course video below is a short introduction to AI/ML with many visuals and I encourage you to watch it if you're interested in the basics.