Europe
DataRobot Looks to Cut Data Science Backlog
The data science automation specialist DataRobot Inc. is gaining traction in the big data market for its machine-learning application as new investors like Intel Capital fund its expanding operations. Boston-based DataRobot has so far raised more than 57 million in four equity investment rounds, including a 33 million funding round completed in February. Along with Intel Capital, Recruit Strategic Partners joined the startup's fourth funding round as new investors. The company's machine-learning platform runs either on top of Hadoop as a cloud application or running on-premises. The startup founded in 2012 by CEO Jerry Achin and CTO Thomas DeGodoy targets its platform at data scientists with varying skill levels. Along with speeding the deployment of more accurate predictive models, the company said it is attempting to address the critical shortage of qualified data scientists while "changing the speed and economics of predictive analytics."
Legal case for drone strikes 'unclear'
The legal case for using drone strikes outside of armed conflict needs "urgent clarification" from ministers, a cross-party parliamentary committee has said. The government insists it does not have a "targeted killing" policy, but the UK was clearly willing to use lethal force overseas for counter-terrorism, the Joint Committee on Human Rights said. It follows the killing of a UK citizen in Syria last year by an RAF drone. The government says it takes "lawful action" over direct threats to the UK. Reyaad Khan, a British member of the so-called Islamic State group, was killed by an RAF drone in Syria last August.
The White House has significant concerns about artificial intelligence
If your mind instantly went to Skynet, I can put your mind at ease; it's not Skynet. That's not, however, to say that this problem isn't just as scary, only without the cool special effects. While Sci-Fi has made the risks of robot takeover well-known, the more immediate concerns are the subtle decisions being made by (sometimes) poorly coded, or designed, algorithms that can drastically alter each of our lives. Our biggest ever edition of TNW Conference is fast approaching! The Obama administration published a report this week that examines problems associated with the shift to an increasingly automated world.
'Viv' is a next-gen AI assistant that runs circles around Siri, Alexa and Cortana
One of the brains behind Apple's iconic digital assistant Siri, Dag Kittlaus, took to the stage today at TechCrunch Disrupt NYC to show off his new project, Viv -- an AI assistant that aims to be "the intelligent interface for everything." "Will it be warmer than 70-degrees near the Golden Gate Bridge after 5pm the day after tomorrow?" Our biggest ever edition of TNW Conference is fast approaching! Viv delivered the answer in a matter of seconds, and correctly handled numerous oddly specific questions that followed, as well as ordering flowers, sending money as payment for drinks the previous night and even booking a hotel in under 60 seconds. Throughout the demo, the AI assistant demonstrated a profound understanding of context and intent as Kittlaus rapid-fired commands in natural language.
An efficient K-means algorithm for Massive Data
Capรณ, Marco, Pรฉrez, Aritz, Lozano, Josรฉ Antonio
Due to the progressive growth of the amount of data available in a wide variety of scientific fields, it has become more difficult to ma- nipulate and analyze such information. Even though datasets have grown in size, the K-means algorithm remains as one of the most popular clustering methods, in spite of its dependency on the initial settings and high computational cost, especially in terms of distance computations. In this work, we propose an efficient approximation to the K-means problem intended for massive data. Our approach recursively partitions the entire dataset into a small number of sub- sets, each of which is characterized by its representative (center of mass) and weight (cardinality), afterwards a weighted version of the K-means algorithm is applied over such local representation, which can drastically reduce the number of distances computed. In addition to some theoretical properties, experimental results indicate that our method outperforms well-known approaches, such as the K-means++ and the minibatch K-means, in terms of the relation between number of distance computations and the quality of the approximation.
Destination Prediction by Trajectory Distribution Based Model
Besse, Philippe C., Guillouet, Brendan, Loubes, Jean-Michel, Royer, Francois
ONITORING and predicting road traffic is of great importance for traffic managers. With the increase of mobile sensors, such as GPS devices and smartphones, much information is at hand to understand urban traffic. In the last few years, a large amount of research has been conducted in order to use this data to model and analyze road traffic conditions. The aim of this paper is to tackle the issue of predicting the destination of vehicles given a prefix of their trajectory. This problem has been the subject of a Kaggle challenge entitled "ECML/PKDD 15: Taxi Trajectory Prediction (I)" [1]. The observations are time-stamped locations that correspond to the different positions of vehicles moving within a city monitored at different observation times. When dealing with a dataset composed of trajectories, the difficulty lies in the fact that the data convey both spatial information (locations of the vehicles on the map of the city) and temporal information (for each vehicle, the locations are indexed by time, which creates a sequence of locations that compose a full trajectory). Hence the data have a spatiotemporal structure that must be taken into account in order to model their evolution while the trajectories of the destination points to be predicted are unknown. Vehicle trajectories are also constrained to a road network which makes their time progression very irregular.
Learning theory estimates with observations from general stationary stochastic processes
Hang, Hanyuan, Feng, Yunlong, Steinwart, Ingo, Suykens, Johan A. K.
This paper investigates the supervised learning problem with observations drawn from certain general stationary stochastic processes. Here by \emph{general}, we mean that many stationary stochastic processes can be included. We show that when the stochastic processes satisfy a generalized Bernstein-type inequality, a unified treatment on analyzing the learning schemes with various mixing processes can be conducted and a sharp oracle inequality for generic regularized empirical risk minimization schemes can be established. The obtained oracle inequality is then applied to derive convergence rates for several learning schemes such as empirical risk minimization (ERM), least squares support vector machines (LS-SVMs) using given generic kernels, and SVMs using Gaussian kernels for both least squares and quantile regression. It turns out that for i.i.d.~processes, our learning rates for ERM recover the optimal rates. On the other hand, for non-i.i.d.~processes including geometrically $\alpha$-mixing Markov processes, geometrically $\alpha$-mixing processes with restricted decay, $\phi$-mixing processes, and (time-reversed) geometrically $\mathcal{C}$-mixing processes, our learning rates for SVMs with Gaussian kernels match, up to some arbitrarily small extra term in the exponent, the optimal rates. For the remaining cases, our rates are at least close to the optimal rates. As a by-product, the assumed generalized Bernstein-type inequality also provides an interpretation of the so-called "effective number of observations" for various mixing processes.
This Week in Fraud & Big Data Technology โ May 6, 2016 โ Fraud & Technology Wire
Here are this week's top stories in fraud and big data technology: Thanks to its huge network of users, Sprint has access to vast amounts of user data. Three years ago it established subsidiary Pinsight Media to investigate ways of capitalizing on that data. Since then it has gone from serving zero to six billion ad impressions per months, based on "authenticated first party data" which it alone has access to. The fact that plain passwords are no longer safe to protect our digital identities is no secret. For years, the use of two-factor authentication (2FA) and multi-factor authentication (MFA) as a means to ensure online account security and prevent fraud has been a hot topic of discussion.
Exploring The Risks Of Artificial Intelligence
"Science has not yet mastered prophecy. We predict too much for the next year and yet far too little for the next ten." These words, articulated by Neil Armstrong at a speech to a joint session of Congress in 1969, fit squarely into most every decade since the turn of the century, and it seems to safe to posit that the rate of change in technology has accelerated to an exponential degree in the last two decades, especially in the areas of artificial intelligence and machine learning. Artificial intelligence is making an extreme entrance into almost every facet of society in predicted and unforeseen ways, causing both excitement and trepidation. This reaction alone is predictable, but can we really predict the associated risks involved?
Picturesqe uses AI to help professional photographers cut through the dross
The word AI is increasingly bandied about with notable abandon. And, as I like to say, AI is the promise that keeps on promising. So I'm always encouraged when I see a humble but viable use of artificial intelligence, even if just a few years ago we may have simply called the nascent technology machine learning. The latest example is Picturesqe, a tool for professional photographers that uses what the U.S./Hungary startup calls AI-powered automation to help pick out the best snaps and filter out the dross. Features of the Windows app and Adobe Lightroom plugin (with a Mac version to follow shortly) include smart grouping, which automatically groups similar photos, intelligent zoom so that you can quickly compare the same spot on multiple shots, and the option to have Picturesqe pick out the best photos for you and delete the duds.