Goto

Collaborating Authors

 Country


New Zealand farmers have a new tool for herding sheep: drones that bark like dogs

#artificialintelligence

You have probably read about robots replacing human labor as a new era of automation takes root in one industry after another. But a new report suggests humans are not the only ones who might lose their jobs. In New Zealand, farmers are using drones to herd and monitor livestock, assuming a job that highly intelligent dogs have held for more than a century. The robots have not replaced the dogs entirely, Radio New Zealand reports, but they have appropriated one of the animal's most potent tools: barking. The DJI Mavic Enterprise, a $3,500 drone favored by farmers, has a feature that lets the machine record sounds and play them over a loud speaker, giving the machine the ability to mimic its canine counterparts.


Age of AI -- The Paradigm Shift to Natural UI

#artificialintelligence

I always loved products and technology. But ever since I was a child, I was especially fascinated by these big inventions, powered by transformative technological revolution that changed - everything! So I felt extremely lucky, when about 20 years ago, at the beginning of my career, I was just in time for one of these revolutions: when the Internet happened. Through the connected PC, the world we lived in has been transformed from a "physical world" -- where we used to go to places like libraries, and use things like encyclopedias and paper maps, to a "digital world" -- where we consume digital information and services from the convenience of our home. What was especially amazing, was the rate and scale of this transformation. Within a very short time, hundreds of millions of users were connected to hundreds of millions online services.


Quantum computing should supercharge this machine-learning technique

#artificialintelligence

Quantum computing and artificial intelligence are both hyped ridiculously. But it seems a combination of the two may indeed combine to open up new possibilities. In a research paper published today in the journal Nature, researchers from IBM and MIT show how an IBM quantum computer can accelerate a specific type of machine-learning task called feature matching. The team says that future quantum computers should allow machine learning to hit new levels of complexity. As first imagined decades ago, quantum computers were seen as a different way to compute information.


The Future Of Recruiting In The Age of Automation And Artificial Intelligence

#artificialintelligence

Despite working in the human resources for over eight years for some of the top organizations in Canada, I have been somewhat blind to the future of work being driven by automation and artificial intelligence (AI). That is, until I read the 2018 Future of Jobs report by World Economic Forum. According to that report, 50% of companies expect that by 2022, their full-time workforce will be somewhat reduced by automation. It is alarming to note that nearly a quarter of these companies are undecided or unlikely to retrain their existing employees and two-thirds expect workers to retrain themselves. This expectation raises an important question of how employees could know which new jobs they should retrain for when these jobs have not yet been created.


PolyAI scores $12M Series A to put its 'conversational AI agents' in contact centres

#artificialintelligence

PolyAI, a London startup founded by experts in the field of "conversational AI" -- including CEO Nikola Mrkลกiฤ‡, who was previously the first engineer at Apple-acquired VocalIQ -- has raised $12 million in Series A funding to deploy its tech in customer support contact centres. The round was led by Point72 Ventures, with participation from Sands Capital Ventures, Amadeus Capital Partners, Passion Capital and Entrepreneur First (EF). PolyAI's founders are graduates of EF, although they didn't meet during the company building program but already knew each other from their time at Cambridge's Dialog Systems Group, part of the Machine Intelligence Lab at the University of Cambridge. "We started PolyAI in 2017, straight after submitting our PhD theses," Mrkลกiฤ‡ tells me. "At Cambridge, we developed state-of-the-art conversational technology, and starting a company was the best way to get this tech used in the real world. We brought many of our Cambridge colleagues with us and started building the commercial version of our conversational platform."


Training Over-parameterized Deep ResNet Is almost as Easy as Training a Two-layer Network

arXiv.org Machine Learning

Although deep neural networks have achieved revolutionary success over various tasks, i.e., computer vision [He et al., 2016] and natural language understanding [Hochreiter and Schmidhuber, 1997], they are still in lack of a rigorous theoretical study of the optimization and generalization properties. Specifically for the optimization, because the loss of deep neural network is highly nonconvex, local search algorithms like gradient descent is hard to analyze with performance guarantee. Many recent works [Choromanska et al., 2015, Kawaguchi, 2016, Nguyen and Hein, 2017, Soudry and Hoffer, 2017] have studied the loss surface of the neural networks and a common claim is that (deep) neural networks have H. Zhang, W. Chen and TY Liu are with Microsoft Research Asia, Beijing, 100080 China (email: {huzhang, wche, tyliu}@microsoft.com); D. Yu is with School of Data and Computer Science at Sun Yat-sen University, Guangzhou, 510275, China (email: yuda3@mail2.sysu.edu.cn).


Training recurrent neural networks robust to incomplete data: application to Alzheimer's disease progression modeling

arXiv.org Machine Learning

Disease progression modeling (DPM) using longitudinal data is a challenging machine learning task. Existing DPM algorithms neglect temporal dependencies among measurements, make parametric assumptions about biomarker trajectories, do not model multiple biomarkers jointly, and need an alignment of subjects' trajectories. In this paper, recurrent neural networks (RNNs) are utilized to address these issues. However, in many cases, longitudinal cohorts contain incomplete data, which hinders the application of standard RNNs and requires a pre-processing step such as imputation of the missing values. Instead, we propose a generalized training rule for the most widely used RNN architecture, long short-term memory (LSTM) networks, that can handle both missing predictor and target values. The proposed LSTM algorithm is applied to model the progression of Alzheimer's disease (AD) using six volumetric magnetic resonance imaging (MRI) biomarkers, i.e., volumes of ventricles, hippocampus, whole brain, fusiform, middle temporal gyrus, and entorhinal cortex, and it is compared to standard LSTM networks with data imputation and a parametric, regression-based DPM method. The results show that the proposed algorithm achieves a significantly lower mean absolute error (MAE) than the alternatives with p < 0.05 using Wilcoxon signed rank test in predicting values of almost all of the MRI biomarkers. Moreover, a linear discriminant analysis (LDA) classifier applied to the predicted biomarker values produces a significantly larger AUC of 0.90 vs. at most 0.84 with p < 0.001 using McNemar's test for clinical diagnosis of AD. Inspection of MAE curves as a function of the amount of missing data reveals that the proposed LSTM algorithm achieves the best performance up until more than 74% missing values. Finally, it is illustrated how the method can successfully be applied to data with varying time intervals.


DSPG: Decentralized Simultaneous Perturbations Gradient Descent Scheme

arXiv.org Machine Learning

In this paper, we present an asynchronous approximate gradient method that is easy to implement called DSPG (Decentralized Simultaneous Perturbation Stochastic Approximations, with Constant Sensitivity Parameters). It is obtained by modifying SPSA (Simultaneous Perturbation Stochastic Approximations) to allow for decentralized optimization in multi-agent learning and distributed control scenarios. SPSA is a popular approximate gradient method developed by Spall, that is used in Robotics and Learning. In the multi-agent learning setup considered herein, the agents are assumed to be asynchronous (agents abide by their local clocks) and communicate via a wireless medium, that is prone to losses and delays. We analyze the gradient estimation bias that arises from setting the sensitivity parameters to a single value, and the bias that arises from communication losses and delays. Specifically, we show that these biases can be countered through better and frequent communication and/or by choosing a small fixed value for the sensitivity parameters. We also discuss the variance of the gradient estimator and its effect on the rate of convergence. Finally, we present numerical results supporting DSPG and the aforementioned theories and discussions.


Learning Competitive and Discriminative Reconstructions for Anomaly Detection

arXiv.org Machine Learning

Most of the existing methods for anomaly detection use only positive data to learn the data distribution, thus they usually need a pre-defined threshold at the detection stage to determine whether a test instance is an outlier. Unfortunately, a good threshold is vital for the performance and it is really hard to find an optimal one. In this paper, we take the discriminative information implied in unlabeled data into consideration and propose a new method for anomaly detection that can learn the labels of unlabelled data directly. Our proposed method has an end-to-end architecture with one encoder and two decoders that are trained to model inliers and outliers' data distributions in a competitive way. This architecture works in a discriminative manner without suffering from overfitting, and the training algorithm of our model is adopted from SGD, thus it is efficient and scalable even for large-scale datasets. Empirical studies on 7 datasets including KDD99, MNIST, Caltech-256, and ImageNet etc. show that our model outperforms the state-of-the-art methods.


Model-Free Model Reconciliation

arXiv.org Artificial Intelligence

Designing agents capable of explaining complex sequential decisions remain a significant open problem in automated decision-making. Recently, there has been a lot of interest in developing approaches for generating such explanations for various decision-making paradigms. One such approach has been the idea of {\em explanation as model-reconciliation}. The framework hypothesizes that one of the common reasons for the user's confusion could be the mismatch between the user's model of the task and the one used by the system to generate the decisions. While this is a general framework, most works that have been explicitly built on this explanatory philosophy have focused on settings where the model of user's knowledge is available in a declarative form. Our goal in this paper is to adapt the model reconciliation approach to the cases where such user models are no longer explicitly provided. We present a simple and easy to learn labeling model that can help an explainer decide what information could help achieve model reconciliation between the user and the agent.