Genre
Adaptive Concentration of Regression Trees, with Application to Random Forests
Wager, Stefan, Walther, Guenther
We study the convergence of the predictive surface of regression trees and forests. To support our analysis we introduce a notion of adaptive concentration for regression trees. This approach breaks tree training into a model selection phase in which we pick the tree splits, followed by a model fitting phase where we find the best regression model consistent with these splits. We then show that the fitted regression tree concentrates around the optimal predictor with the same splits: as d and n get large, the discrepancy is with high probability bounded on the order of sqrt(log(d) log(n)/k) uniformly over the whole regression surface, where d is the dimension of the feature space, n is the number of training examples, and k is the minimum leaf size for each tree. We also provide rate-matching lower bounds for this adaptive concentration statement. From a practical perspective, our result enables us to prove consistency results for adaptively grown forests in high dimensions, and to carry out valid post-selection inference in the sense of Berk et al. [2013] for subgroups defined by tree leaves.
Computational Cost Reduction in Learned Transform Classifications
Machado, Emerson Lopes, Miosso, Cristiano Jacques, von Borries, Ricardo, Coutinho, Murilo, Berger, Pedro de Azevedo, Marques, Thiago, Jacobi, Ricardo Pezzuol
We present a theoretical analysis and empirical evaluations of a novel set of techniques for computational cost reduction of classifiers that are based on learned transform and soft-threshold. By modifying optimization procedures for dictionary and classifier training, as well as the resulting dictionary entries, our techniques allow to reduce the bit precision and to replace each floating-point multiplication by a single integer bit shift. We also show how the optimization algorithms in some dictionary training methods can be modified to penalize higher-energy dictionaries. We applied our techniques with the classifier Learning Algorithm for Soft-Thresholding, testing on the datasets used in its original paper. Our results indicate it is feasible to use solely sums and bit shifts of integers to classify at test time with a limited reduction of the classification accuracy. These low power operations are a valuable trade off in FPGA implementations as they increase the classification throughput while decrease both energy consumption and manufacturing cost.
An Improved System for Sentence-level Novelty Detection in Textual Streams
Fu, Xinyu, Ch'ng, Eugene, Aickelin, Uwe, Zhang, Lanyun
Novelty detection in news events has long been a difficult problem. A number of models performed well on specific data streams but certain issues are far from being solved, particularly in large data streams from the WWW where unpredictability of new terms requires adaptation in the vector space model. We present a novel event detection system based on the Incremental Term Frequency-Inverse Document Frequency (TF-IDF) weighting incorporated with Locality Sensitive Hashing (LSH). Our system could efficiently and effectively adapt to the changes within the data streams of any new terms with continual updates to the vector space model. Regarding miss probability, our proposed novelty detection framework outperforms a recognised baseline system by approximately 16% when evaluating a benchmark dataset from Google News.
Unsupervised and Semi-supervised Learning with Categorical Generative Adversarial Networks
In this paper we present a method for learning a discriminative classifier from unlabeled or partially labeled data. Our approach is based on an objective function that trades-off mutual information between observed examples and their predicted categorical class distribution, against robustness of the classifier to an adversarial generative model. The resulting algorithm can either be interpreted as a natural generalization of the generative adversarial networks (GAN) framework or as an extension of the regularized information maximization (RIM) framework to robust classification against an optimal adversary. We empirically evaluate our method - which we dub categorical generative adversarial networks (or CatGAN) - on synthetic data as well as on challenging image classification tasks, demonstrating the robustness of the learned classifiers. We further qualitatively assess the fidelity of samples generated by the adversarial generator that is learned alongside the discriminative classifier, and identify links between the CatGAN objective and discriminative clustering algorithms (such as RIM).
The wonderful world of recommender systems
I recently gave a talk about recommender systems at the Data Science Sydney meetup (the slides are available here). This post roughly follows the outline of the talk, expanding on some of the key points in non-slide form (i.e., complete sentences and paragraphs!). The first few sections give a broad overview of the field and the common recommendation paradigms, while the final part is dedicated to debunking five common myths about recommender systems. The key reason why many people seem to care about recommender systems is money. For companies such as Amazon, Netflix, and Spotify, recommender systems drive significant engagement and revenue. But this is the more cynical view of things.
'Machine learning' may contribute to new advances in plastic surgery
April 29, 2016 - With an ever-increasing volume of electronic data being collected by the healthcare system, researchers are exploring the use of machine learning--a subfield of artificial intelligence--to improve medical care and patient outcomes. An overview of machine learning and some of the ways it could contribute to advancements in plastic surgery are presented in a special topic article in the May issue of Plastic and Reconstructive Surgery, the official medical journal of the American Society of Plastic Surgeons (ASPS). "Machine learning has the potential to become a powerful tool in plastic surgery, allowing surgeons to harness complex clinical data to help guide key clinical decision-making," write Dr. Jonathan Kanevsky of McGill University, Montreal, and colleagues. They highlight some key areas in which machine learning and "Big Data" could contribute to progress in plastic and reconstructive surgery. Machine learning analyzes historical data to develop algorithms capable of knowledge acquisition.
Mark Zuckerberg Thinks Elon Musk is Wrong on AI -- The Motley Fool
Tesla CEO Elon Musk introduces the Model X. I think we should be very careful about artificial intelligence. If I had to guess at what our biggest existential threat is, it's probably that. So we need to be very careful with artificial intelligence...With artificial intelligence, we're summoning the demon. You know those stories where there's the guy with the pentagram, and the holy water, and he's like -- Yeah, he's sure he can control the demon?
Apple Shows Us It's Hard to Be Innovative When You're on Top. But Does it Really Matter? Fox News
Once your business is no longer the innovative upstart and you become the establishment entity, how do you maintain an entrepreneurial and disruptive spirit that gets results? That's the question Apple had to ask itself this week, following an iffy earnings report. This week, Apple posted the earnings results for the second quarter of 2016, and reported a year-over-year decline in quarterly revenue for the first time in 13 years. The company took in 50.6 billion in quarterly revenue and 10.5 billion in quarterly net income. On a call with investors, CEO Tim Cook characterized that 13 percent dip in revenue as a "pause in our growth," that had stemmed from "ongoing macroeconomic headwinds in much of the world." Despite the break in the company's decade plus streak of "record" growth, it's unlikely that the tech giant's standing as one of most valuable and authentic brands in the world will be dinged in any significant way.
Apple Shows Us It's Hard to Be Innovative When You're on Top. But Does it Really Matter?
Apply now to be an Enterpreneur360 company and let us tell the world your success story. Once your business is no longer the innovative upstart and you become the establishment entity, how do you maintain an entrepreneurial and disruptive spirit that gets results? That's the question Apple had to ask itself this week, following an iffy earnings report. This week, Apple posted the earnings results for the second quarter of 2016, and reported a year-over-year decline in quarterly revenue for the first time in 13 years. The company took in 50.6 billion in quarterly revenue and 10.5 billion in quarterly net income. On a call with investors, CEO Tim Cook characterized that 13 percent dip in revenue as a "pause in our growth," that had stemmed from "ongoing macroeconomic headwinds in much of the world."
Newly declassified pictures show USS Independence as it was blown up alongside 77 other ships as part of atomic tests at Bikini Atoll in 1946
Stunning new pictures from a 1946 atomic weapon test on a hundred US ships have been revealed. The newly declassified images show the World War II veteran aircraft carrier USS Independence, which was one of nearly a hundred ships used as targets in the first tests of the atomic bomb at Bikini Atoll in the summer of 1946. The two Bikini tests known as Operation Crossroads were carried out in the immediate aftermath of the atomic end to World War II in Japan, and signaled a new era in world history, the historians involved in the new study say. The newly declassified images show the World War II aircraft carrier which was one of nearly a hundred ships used as targets in the first tests of the atomic bomb at Bikini Atoll in 1946. Here, Sailors watch the'Able Test' burst miles out to sea from the deck of the support ship USS Fall River on 1 July 1946.