Genre
Chris Dixon on competing with Internet giants for budding AI and VR talent
VC Chris Dixon of Andreessen Horowitz thinks it's a lot harder to predict financial cycles than it is to see a new computing platform coming down the pike. As he noted in a recent post, new cycles tend to begin every 10 to 15 years; assuming the 2007 introduction of the iPhone kicked off the last wave, we're fast heading toward the Next New Thing. Or things, technically, according to Dixon, who we caught up with yesterday. Among the trends that Dixon is watching closely, he says, are virtual reality, augmented reality, IoT, wearables, drones and cars. Not that it'll be easy to make money off these newer technologies. In fact, Dixon suggests it could be ridiculously challenging, given how quickly Facebook, Google, and Amazon are bringing aboard related talent.
Data Science Has Been Using Rebel Statistics for a Long Time
Many of those who call themselves statisticians just won't admit that data science heavily relies on and uses (heretical, rule-breaking) statistical science, or they don't recognize the true statistical nature of these data science techniques (some are 15-year old), or are opposed to the modernization of their statistical arsenal. They already missed the train when machine learning became a popular discipline (also heavily based on statistics) more than 15 years ago. Now machine learning professionals, who are statistical practitioners working on problems such as clustering, far outnumber statisticians. Many times, I have interacted with statisticians who think that anyone not calling himself statistician, knows nothing or little about statistics; see my recent bio published here, or visit the LinkedIn profiles of many data scientists, to debunk this myth. Any statistical technique that is not in their old books are considered heretical at best, or non-statistic at worst, or most of the time, not understood.
Semiconductor Engineering .:. System Bits: April 19
Debugging web apps MIT researchers reported that they've developed a system that can quickly comb through tens of thousands of lines of application code to find security flaws by exploiting some peculiarities of the Ruby on Rails web programming framework. The team said that in tests on 50 popular web applications written using Ruby on Rails, the system found 23 previously undiagnosed security flaws, and it took no more than 64 seconds to analyze any given program. Daniel Jackson, professor in the Department of Electrical Engineering and Computer Science, said the system uses static analysis, which seeks to describe, in a very general way, how data flows through a program. "The classic example of this is if you wanted to do an abstract analysis of a program that manipulates integers, you might divide the integers into the positive integers, the negative integers, and zero." The static analysis would then evaluate every operation in the program according to its effect on integers' signs.
Human eyes only see the gist of scenes before us and miss the details, research suggests
They are said to the be windows to the soul, but our eyes may work more like shutters if recent scientific research is to be believed Some neuroscientists argue regardless of what details a person notes, the eyes still captured everything in front of them much like a camera. But a team of researchers has concluded that our eyes in fact only reflect the gist of the scenes before us.
A Factorization Machine Framework for Testing Bigram Embeddings in Knowledgebase Completion
Welbl, Johannes, Bouchard, Guillaume, Riedel, Sebastian
Embedding-based Knowledge Base Completion models have so far mostly combined distributed representations of individual entities or relations to compute truth scores of missing links. Facts can however also be represented using pairwise embeddings, i.e. embeddings for pairs of entities and relations. In this paper we explore such bigram embeddings with a flexible Factorization Machine model and several ablations from it. We investigate the relevance of various bigram types on the fb15k237 dataset and find relative improvements compared to a compositional model.
Random Projection Estimation of Discrete-Choice Models with Large Choice Sets
Chiong, Khai X., Shum, Matthew
We introduce sparse random projection, an important dimension-reduction tool from machine learning, for the estimation of discrete-choice models with high-dimensional choice sets. Initially, high-dimensional data are compressed into a lower-dimensional Euclidean space using random projections. Subsequently, estimation proceeds using cyclic monotonicity moment inequalities implied by the multinomial choice model; the estimation procedure is semi-parametric and does not require explicit distributional assumptions to be made regarding the random utility errors. The random projection procedure is justified via the Johnson-Lindenstrauss Lemma -- the pairwise distances between data points are preserved during data compression, which we exploit to show convergence of our estimator. The estimator works well in simulations and in an application to a supermarket scanner dataset.
Constructive Preference Elicitation by Setwise Max-margin Learning
Teso, Stefano, Passerini, Andrea, Viappiani, Paolo
In this paper we propose an approach to preference elicitation that is suitable to large configuration spaces beyond the reach of existing state-of-the-art approaches. Our setwise max-margin method can be viewed as a generalization of max-margin learning to sets, and can produce a set of "diverse" items that can be used to ask informative queries to the user. Moreover, the approach can encourage sparsity in the parameter space, in order to favor the assessment of utility towards combinations of weights that concentrate on just few features. We present a mixed integer linear programming formulation and show how our approach compares favourably with Bayesian preference elicitation alternatives and easily scales to realistic datasets.
Trading-Off Cost of Deployment Versus Accuracy in Learning Predictive Models
Robinson, Daniel P., Saria, Suchi
Predictive models are finding an increasing number of applications in many industries. As a result, a practical means for trading-off the cost of deploying a model versus its effectiveness is needed. Our work is motivated by risk prediction problems in healthcare. Cost-structures in domains such as healthcare are quite complex, posing a significant challenge to existing approaches. We propose a novel framework for designing cost-sensitive structured regularizers that is suitable for problems with complex cost dependencies. We draw upon a surprising connection to boolean circuits. In particular, we represent the problem costs as a multi-layer boolean circuit, and then use properties of boolean circuits to define an extended feature vector and a group regularizer that exactly captures the underlying cost structure. The resulting regularizer may then be combined with a fidelity function to perform model prediction, for example. For the challenging real-world application of risk prediction for sepsis in intensive care units, the use of our regularizer leads to models that are in harmony with the underlying cost structure and thus provide an excellent prediction accuracy versus cost tradeoff.
Alphabet Inc (GOOG) Q1 2016 Earnings Preview: Big Profits Despite EU Challenges, Unprofitable Moonshots
It's a good time to be Alphabet Inc. (GOOG), the parent company of Google. The holding company that owns Google, YouTube and Android -- as well as so-called moonshots like self-driving cars, the home-networking division Nest and Google Fiber -- is expected to turn in healthy first-quarter results on Thursday, driven by its dominant position in online search and display advertising. On Wednesday, the European Commission is expected to formally charge Google for favoring its own apps and services on its Android mobile operating system, which powers more than 80 percent of the world's smartphones. That will be the latest in a decade of entanglements with regulators on both sides of the Atlantic; Google also got some bad press in Britain earlier this year for having paid just 185 million in taxes over the past decade. Also confronting Google -- and the rest of the tech industry -- is how to manage government and law enforcement requests for information.