Goto

Collaborating Authors

 Europe


Spatial Decompositions for Large Scale SVMs

arXiv.org Machine Learning

Although support vector machines (SVMs) are theoretically well understood, their underlying optimization problem becomes very expensive, if, for example, hundreds of thousands of samples and a non-linear kernel are considered. Several approaches have been proposed in the past to address this serious limitation. In this work we investigate a decomposition strategy that learns on small, spatially defined data chunks. Our contributions are two fold: On the theoretical side we establish an oracle inequality for the overall learning method using the hinge loss, and show that the resulting rates match those known for SVMs solving the complete optimization problem with Gaussian kernels. On the practical side we compare our approach to learning SVMs on small, randomly chosen chunks. Here it turns out that for comparable training times our approach is significantly faster during testing and also reduces the test error in most cases significantly. Furthermore, we show that our approach easily scales up to 10 million training samples: including hyper-parameter selection using cross validation, the entire training only takes a few hours on a single machine. Finally, we report an experiment on 32 million training samples. All experiments used liquidSVM (Steinwart and Thomann, 2017).


Go With the Flow, on Jupiter and Snow. Coherence From Model-Free Video Data without Trajectories

arXiv.org Machine Learning

Viewing a data set such as the clouds of Jupiter, coherence is readily apparent to human observers, especially the Great Red Spot, but also other great storms and persistent structures. There are now many different definitions and perspectives mathematically describing coherent structures, but we will take an image processing perspective here. We describe an image processing perspective inference of coherent sets from a fluidic system directly from image data, without attempting to first model underlying flow fields, related to a concept in image processing called motion tracking. In contrast to standard spectral methods for image processing which are generally related to a symmetric affinity matrix, leading to standard spectral graph theory, we need a not symmetric affinity which arises naturally from the underlying arrow of time. We develop an anisotropic, directed diffusion operator corresponding to flow on a directed graph, from a directed affinity matrix developed with coherence in mind, and corresponding spectral graph theory from the graph Laplacian. Our methodology is not offered as more accurate than other traditional methods of finding coherent sets, but rather our approach works with alternative kinds of data sets, in the absence of vector field. Our examples will include partitioning the weather and cloud structures of Jupiter, and a local to Potsdam, N.Y. lake-effect snow event on Earth, as well as the benchmark test double-gyre system.


Optimization Methods for Large-Scale Machine Learning

arXiv.org Machine Learning

This paper provides a review and commentary on the past, present, and future of numerical optimization algorithms in the context of machine learning applications. Through case studies on text classification and the training of deep neural networks, we discuss how optimization problems arise in machine learning and what makes them challenging. A major theme of our study is that large-scale machine learning represents a distinctive setting in which the stochastic gradient (SG) method has traditionally played a central role while conventional gradient-based nonlinear optimization techniques typically falter. Based on this viewpoint, we present a comprehensive theory of a straightforward, yet versatile SG algorithm, discuss its practical behavior, and highlight opportunities for designing algorithms with improved performance. This leads to a discussion about the next generation of optimization methods for large-scale machine learning, including an investigation of two main streams of research on techniques that diminish noise in the stochastic directions and methods that make use of second-order derivative approximations.


Detection of Adversarial Training Examples in Poisoning Attacks through Anomaly Detection

arXiv.org Machine Learning

Machine learning has become an important component for many systems and applications including computer vision, spam filtering, malware and network intrusion detection, among others. Despite the capabilities of machine learning algorithms to extract valuable information from data and produce accurate predictions, it has been shown that these algorithms are vulnerable to attacks. Data poisoning is one of the most relevant security threats against machine learning systems, where attackers can subvert the learning process by injecting malicious samples in the training data. Recent work in adversarial machine learning has shown that the so-called optimal attack strategies can successfully poison linear classifiers, degrading the performance of the system dramatically after compromising a small fraction of the training dataset. In this paper we propose a defence mechanism to mitigate the effect of these optimal poisoning attacks based on outlier detection. We show empirically that the adversarial examples generated by these attack strategies are quite different from genuine points, as no detectability constrains are considered to craft the attack. Hence, they can be detected with an appropriate pre-filtering of the training dataset.


Future risks associated with machine learning explored in new report

#artificialintelligence

A new study released by The Economist Intelligence Unit ran three econometric scenarios to 2030 on five countries -- the United States, the United Kingdom, Australia, Japan--and developing Asia as a whole. In'Risks and rewards: Scenarios around the economic impact of machine learning', commissioned by Google, two scenarios assumed greater human productivity through upskilling and greater investment in technology and access to open source data, while the third assumed insufficient policy support for structural changes in the economy. The results showed that, although the fears of those pessimistic about the impact of machine learning, and artificial intelligence in general, may be overblown, the optimists' claims are not entirely supported, either. The other area of the study, a look at the impact of machine learning on four industries, reaches a similar conclusion. For firms both developing machine learning and those using it, the reports finds that communication between themselves, and with the public and policymakers, needs to improve.


Artificial Intelligence in Transportation Industry Is Moving Fast Here Is What You Need To Know

#artificialintelligence

The artificial intelligence in transportation market is projected to grow at a CAGR of 17.87% from 2017 to 2030, and the market size is expected to grow from USD 1.21 Billion in 2017 to USD 10.30 Billion by 2030. The increasing government regulations for vehicle safety, growing adoption of advanced driver assistance systems (ADAS), and development of autonomous vehicles play a significant role in the growth of this market. Deep learning technology is estimated to be the largest and fastest growing segment of the artificial intelligence in transportation market, by machine learning technology. The deep learning technology is widely used in the development of autonomous vehicles, which need to see, think, drive, and learn. The last step is "learn," where deep learning will be critical for achieving fully autonomous vehicles.


Finally The Secret to Risk Modeling with Python Provenir

#artificialintelligence

It was the week of Christmas in 1989 in Amsterdam, 48* F and cloudy, when Guido van Rossum started tinkering on a small project to busy himself while his employer's office closed for the holiday. He placed a high emphasis on readability and uniformity in an era that praised languages like C and Perl (PHP's high-maintenance personality would crash the party later). The result was a gorgeously elegant open source language named after Monty Python's Flying Circus. Since its official release in 1993, Python has displayed its prowess across multiple industries and in widely varying use cases. It is now the most widely taught introductory language in the top computer science programs worldwide, was named the most in-demand programming language in the U.S. by Forbes' fintech columnist, and hit #2 on the list of most GitHub pulls by language in 2017.


Impact of AI and ML on trading and investing

#artificialintelligence

Very few can ignore the presence of Artificial Intelligence and Machine Learning in today's world, and even less so if you work in quantitative finance. Here, Michael Harris, quant systematic and discretionary trader and best selling author, discusses the impact these technologies are having on trading and investing. Below are excerpts from a presentation I gave last year in Europe, as an invited speaker to a group of low profile but high net worth investors and traders. The subject was determined by the organizer to be about the impact of artificial intelligence and machine learning on trading and investing. The excerpts below are organized in four sections and cover about 50% of the original presentation.


The Scientific Alliance

#artificialintelligence

A variety of headlines appear in the Scottish papers this morning, including new research on beating cancer and how the PM alleges that bullying on social media is a threat to democracy. The UK front pages cover calls for a pardon for suffragettes and reaction to Trump's comment on the NHS among other issues.


Artificial intelligence in social housing: your virtual assistant

#artificialintelligence

The rise of artificial intelligence (AI), as with a lot of recent technologies, seems to be exponential. Recently, the news was of the NHS looking to introduce AI to assist in more rapid diagnosis of heart conditions. So perhaps the time is right to start thinking how AI could be utilised within UK housing organisations? At the beginning of 2017 I began an open conversation with the housing sector around emerging technologies, including chatbots, sensor-based Internet of Things devices and headless user interfaces (UIs) like Amazon's Alexa or Google at Home. Since then I've been visiting and meeting with a diverse range of organisations who have all started to explore the potential benefits that these sorts of technologies could offer them.