Genre
Compressive K-means
Keriven, Nicolas, Tremblay, Nicolas, Traonmilin, Yann, Gribonval, Rémi
The Lloyd-Max algorithm is a classical approach to perform K-means clustering. Unfortunately, its cost becomes prohibitive as the training dataset grows large. We propose a compressive version of K-means (CKM), that estimates cluster centers from a sketch, i.e. from a drastically compressed representation of the training dataset. We demonstrate empirically that CKM performs similarly to Lloyd-Max, for a sketch size proportional to the number of cen-troids times the ambient dimension, and independent of the size of the original dataset. Given the sketch, the computational complexity of CKM is also independent of the size of the dataset. Unlike Lloyd-Max which requires several replicates, we further demonstrate that CKM is almost insensitive to initialization. For a large dataset of 10^7 data points, we show that CKM can run two orders of magnitude faster than five replicates of Lloyd-Max, with similar clustering performance on artificial data. Finally, CKM achieves lower classification errors on handwritten digits classification.
Adversarial examples in the physical world
Kurakin, Alexey, Goodfellow, Ian, Bengio, Samy
Most existing machine learning classifiers are highly vulnerable to adversarial examples. An adversarial example is a sample of input data which has been modified very slightly in a way that is intended to cause a machine learning classifier to misclassify it. In many cases, these modifications can be so subtle that a human observer does not even notice the modification at all, yet the classifier still makes a mistake. Adversarial examples pose security concerns because they could be used to perform an attack on machine learning systems, even if the adversary has no access to the underlying model. Up to now, all previous work have assumed a threat model in which the adversary can feed data directly into the machine learning classifier. This is not always the case for systems operating in the physical world, for example those which are using signals from cameras and other sensors as an input. This paper shows that even in such physical world scenarios, machine learning systems are vulnerable to adversarial examples. We demonstrate this by feeding adversarial images obtained from cell-phone camera to an ImageNet Inception classifier and measuring the classification accuracy of the system. We find that a large fraction of adversarial examples are classified incorrectly even when perceived through the camera.
Learning what matters - Sampling interesting patterns
Dzyuba, Vladimir, van Leeuwen, Matthijs
In the field of exploratory data mining, local structure in data can be described by patterns and discovered by mining algorithms. Although many solutions have been proposed to address the redundancy problems in pattern mining, most of them either provide succinct pattern sets or take the interests of the user into account-but not both. Consequently, the analyst has to invest substantial effort in identifying those patterns that are relevant to her specific interests and goals. To address this problem, we propose a novel approach that combines pattern sampling with interactive data mining. In particular, we introduce the LetSIP algorithm, which builds upon recent advances in 1) weighted sampling in SAT and 2) learning to rank in interactive pattern mining. Specifically, it exploits user feedback to directly learn the parameters of the sampling distribution that represents the user's interests. We compare the performance of the proposed algorithm to the state-of-the-art in interactive pattern mining by emulating the interests of a user. The resulting system allows efficient and interleaved learning and sampling, thus user-specific anytime data exploration. Finally, LetSIP demonstrates favourable trade-offs concerning both quality-diversity and exploitation-exploration when compared to existing methods.
$L_2$Boosting for Economic Applications
In the recent years more and more high-dimensional data sets, where the number of parameters $p$ is high compared to the number of observations $n$ or even larger, are available for applied researchers. Boosting algorithms represent one of the major advances in machine learning and statistics in recent years and are suitable for the analysis of such data sets. While Lasso has been applied very successfully for high-dimensional data sets in Economics, boosting has been underutilized in this field, although it has been proven very powerful in fields like Biostatistics and Pattern Recognition. We attribute this to missing theoretical results for boosting. The goal of this paper is to fill this gap and show that boosting is a competitive method for inference of a treatment effect or instrumental variable (IV) estimation in a high-dimensional setting. First, we present the $L_2$Boosting with componentwise least squares algorithm and variants which are tailored for regression problems which are the workhorse for most Econometric problems. Then we show how $L_2$Boosting can be used for estimation of treatment effects and IV estimation. We highlight the methods and illustrate them with simulations and empirical examples. For further results and technical details we refer to Luo and Spindler (2016, 2017) and to the online supplement of the paper.
The internet works like a human BRAIN, a new study finds
A similar rule regulates traffic flow in both the internet and the human brain, a new study has found. Researcher have discovered that our brain has a neuronal equivalent of a flow-control algorithm that checks the internet for congestion. The team believes these findings could improve the understanding of engineered and neural networks and could lead to treatments for learning disabilities. Researcher have discovered that our brain has a neuronal equivalent of a flow-control algorithm that checks the internet for congestion. The algorithm AIMD, sends a packet of data through different routes and then'listen' for confirmation from the receiver An algorithm called'additive increase, multiplicative decrease' (AIMD) checks how congested the internet is.
Nvidia posts record Q4 results but shares fall after hours ZDNet
Nvidia posted its fourth quarter results on Thursday, once again surpassing market expectations. However, its outlook for the current quarter is only slightly above market market consensus. Shares fell in after-hours trading. Non-GAAP earnings for the quarter were $1.13, up 117 percent from a year earlier. Revenue came to $2.17 billion, up 55 percent year-over-year.
Should we replace honeybees with pollinating drones?
February 9, 2017 --Three-quarters of the world's food crops require pollination, according to the Food and Agriculture Organization of the United Nations, but more than 40 percent of the species that perform this vital service are under threat. Researchers across disciplines have been searching for solutions. Some focus on ways to protect the bees and other crucial pollinators. But others are looking outside of the natural world for ways to protect crops like fruits, vegetables, nuts, berries, and even chocolate and coffee. Perhaps an army of robotic pollinators could keep humans well-supplied in these foods, some engineers have thought.
Science confirms what we already know: It's all in the hips
To find out what people think of lady dancing, you don't need to head to the club. Instead, researchers in the UK outfitted female dancers with motion capture rigs, much like the ones that bring digital movie characters like Gollum or Jar Jar Binks to life. According to science, then, women who swing their hips while moving their legs and thighs independently are rated high on attractiveness. In this study, published Thursday, the authors showed the digitized dance avatars to a group of heterosexual men and women. The basic mannequin-like animations allowed the researchers to control for any secondary variables that might denote hotness like clothing or hairstyle.
DeepMind's AI has learnt to become 'highly aggressive' when it feels like it's going to lose
Artificial intelligence changes the way it behaves based on the environment it is in, much like humans do, according to the latest research from DeepMind . Computer scientists from the Google-owned firm have studied how their AI behaves in social situations by using principles from game theory and social sciences. During the work, they found it is possible for AI to act in an "aggressive manner" when it feels it is going to lose out, but agents will work as a team when there is more to be gained. For the research, the AI was tested on two games: a fruit gathering game and a Wolfpack hunting game. These are both basic, 2D games that used AI characters (known as agents) similar to those used in DeepMind's original work with Atari.
As bee populations dwindle, robot bees may help pick up some of their pollination slack
One day, gardeners might not just hear the buzz of bees among their flowers, but the whirr of robots, too. Scientists in Japan say they've managed to turn an unassuming drone into a remote-controlled pollinator by attaching horsehairs coated with a special, sticky gel to its underbelly. The system, described in the journal Chem, is nowhere near ready to be sent to agricultural fields, but it could help pave the way to developing automated pollination techniques at a time when bee colonies are suffering precipitous declines. In flowering plants, sex often involves a threesome. Flowers looking to get the pollen from their male parts into another bloom's female parts need an envoy to carry it from one to the other.