Goto

Collaborating Authors

 Genre


How to detect spurious correlations, and how to find the real ones

@machinelearnbot

Originally posted on DataSciebceCentral, by Dr. Granville. Click here to read original article and comments. Specifically designed in the context of big data in our research lab, the new and simple strong correlation synthetic metric proposed in this article should be used, whenever you want to check if there is a real association between two variables, especially in large-scale automated data science or machine learning projects. Use this new metric now, to avoid being accused of reckless data science and evenbeing sued for wrongful analytic practice. In this paper, the traditional correlation is referred to as the weak correlation, as it captures only a small part of the association between two variables: weak correlation results in capturing spurious correlations and predictive modeling deficiencies, even with as few as 100 variables.


Alibaba invests in Israeli e-commerce search co Twiggle - Globes English

#artificialintelligence

Israeli startup Twiggle, which is developing next generation e-commerce search technologies, announced today that it secured additional funding from the Alibaba Group as the second tranche of its Series A financing. This follows the announcement in April of a 12.5 million round led by Naspers with participation from YJ Capital, State of Mind Ventures and Sir Ronald Cohen. The funding will be utilized to grow the company's R&D team in Israel and drive the company's global expansion plans. No details were disclosed about the amount Alibaba is investing but "Bloomberg" reported that it is 5-10 million. Twiggle uses advanced techniques in data science, artificial intelligence, machine learning and natural language processing to power the next generation of digital commerce.


How Salesforce Is Betting on Artificial Intelligence

#artificialintelligence

According to MarketandMarkets, the artificial intelligence (AI) market is estimated to grow from 419.7 million in 2014 to 5.05 billion by 2020, growing at a CAGR of 53.65% from 2015 to 2020. The Media and Advertising sector is expected to drive the growth of AI during this period. IBM, Microsoft, and Google are key players in the market, and now Salesforce is trying to make inroads into it. For the first quarter of fiscal 2017, Salesforce's revenue grew 27% over the year to 1.92 billion, above analyst estimate of 1.89 billion. Net income was 38.8 billion or 0.06 per share. Non GAAP EPS was 0.24, beating analyst forecast of 0.25.


Rolling Stone Australia -- The Rise of Intelligent Machines: Part 2

#artificialintelligence

It's a weird feeling, cruising around Silicon Valley in a car driven by no one. I am in the back seat of one of Google's self-driving cars – a converted Lexus SUV with lasers, radar and low-res cameras strapped to the roof and fenders – as it manoeuvres the streets of Mountain View, California, not far from Google's headquarters. I grew up about eight kilometres from here and remember riding around on these same streets on a Schwinn Sting-Ray. Now, I am riding an algorithm, you might say – a mathematical equation, which, written as computer code, controls the Lexus. The car does not feel dangerous, nor does it feel like it is being driven by a human. It rolls to a full stop at stop signs, veers too far away from a delivery van, taps the brakes for no apparent reason as we pass a line of parked cars. I wonder if the flaw is in me, not the car: Is it reacting to something I can't see? The car is capable of detecting the motion of a cat, or a car crossing the street hundreds of metres away in any direction, day or night (snow and fog can be another matter). "It sees much better than a human being," Dmitri Dolgov, the lead software engineer for Google's self-driving-car project, says proudly. He is sitting behind the wheel, his hands on his lap. As we stop at the intersection, waiting for a left turn, I glance over at a laptop in the passenger seat that provides a real-time look at how the car interprets its surroundings. On it, I see a gridlike world of colourful objects – cars, trucks, bicyclists, pedestrians – drifting by in a video-game-like tableau. Each sensor offers a different view – the lasers provide three-dimensional depth, the cameras identify road signs, turn signals, colours and lights. The computer in the back processes all this information in real time, gauging the speed of oncoming traffic, making a judgment about when it is OK to make a left turn. Waiting for the car to make that decision is a spooky moment. I am betting my life that one of the coders who worked on the algorithm for when it's safe to make a left-hand turn in traffic had not had a fight with his girlfriend (or boyfriend) the night before and screwed up the code.


SoftBank to sell 8B in Alibaba stock

USATODAY - Tech Top Stories

IBM Watson is going to Japan via IBM's new alliance with Japanese telecommunication giant SoftBank, on Tuesday, February 10, 2015. SAN FRANCISCO -- Japanese telecommunications giant SoftBank is selling 8 billion in Alibaba stock in order to pay down debt, the company said in a statement Tuesday. SoftBank, which was among the earliest investors in Chinese e-commerce juggernaut Alibaba (BABA), will sell 2 billion in stock back to Alibaba. It will offer another 5 billion in securities that in three years will convert into Alibaba shares. Another 500,000 in stock will be sold to an unnamed wealth fund and 400,000 to the Alibaba Partnership, which controls nomination of the company's directors.


Gene-Gene association for Imaging Genetics Data using Robust Kernel Canonical Correlation Analysis

arXiv.org Machine Learning

In genome-wide interaction studies, to detect gene-gene interactions, most methods are divided into two folds: single nucleotide polymorphisms (SNP) based and gene-based methods. Basically, the methods based on the gene are more effective than the methods based on a single SNP. Recent years, while the kernel canonical correlation analysis (Classical kernel CCA) based U statistic (KCCU) has proposed to detect the nonlinear relationship between genes. To estimate the variance in KCCU, they have used resampling based methods which are highly computationally intensive. In addition, classical kernel CCA is not robust to contaminated data. We, therefore, first discuss robust kernel mean element, the robust kernel covariance, and cross-covariance operators. Second, we propose a method based on influence function to estimate the variance of the KCCU. Third, we propose a nonparametric robust KCCU method based on robust kernel CCA, which is designed for contaminated data and less sensitive to noise than classical kernel CCA. Finally, we investigate the proposed methods to synthesized data and imaging genetic data set. Based on gene ontology and pathway analysis, the synthesized and genetics analysis demonstrate that the proposed robust method shows the superior performance of the state-of-the-art methods.


Temporal Topic Modeling to Assess Associations between News Trends and Infectious Disease Outbreaks

arXiv.org Machine Learning

In retrospective assessments, internet news reports have been shown to capture early reports of unknown infectious disease transmission prior to official laboratory confirmation. In general, media interest and reporting peaks and wanes during the course of an outbreak. In this study, we quantify the extent to which media interest during infectious disease outbreaks is indicative of trends of reported incidence. We introduce an approach that uses supervised temporal topic models to transform large corpora of news articles into temporal topic trends. The key advantages of this approach include, applicability to a wide range of diseases, and ability to capture disease dynamics - including seasonality, abrupt peaks and troughs. We evaluated the method using data from multiple infectious disease outbreaks reported in the United States of America (U.S.), China and India. We noted that temporal topic trends extracted from disease-related news reports successfully captured the dynamics of multiple outbreaks such as whooping cough in U.S. (2012), dengue outbreaks in India (2013) and China (2014). Our observations also suggest that efficient modeling of temporal topic trends using time-series regression techniques can estimate disease case counts with increased precision before official reports by health organizations.


Scaling Submodular Maximization via Pruned Submodularity Graphs

arXiv.org Machine Learning

We propose a new random pruning method (called "submodular sparsification (SS)") to reduce the cost of submodular maximization. The pruning is applied via a "submodularity graph" over the $n$ ground elements, where each directed edge is associated with a pairwise dependency defined by the submodular function. In each step, SS prunes a $1-1/\sqrt{c}$ (for $c>1$) fraction of the nodes using weights on edges computed based on only a small number ($O(\log n)$) of randomly sampled nodes. The algorithm requires $\log_{\sqrt{c}}n$ steps with a small and highly parallelizable per-step computation. An accuracy-speed tradeoff parameter $c$, set as $c = 8$, leads to a fast shrink rate $\sqrt{2}/4$ and small iteration complexity $\log_{2\sqrt{2}}n$. Analysis shows that w.h.p., the greedy algorithm on the pruned set of size $O(\log^2 n)$ can achieve a guarantee similar to that of processing the original dataset. In news and video summarization tasks, SS is able to substantially reduce both computational costs and memory usage, while maintaining (or even slightly exceeding) the quality of the original (and much more costly) greedy algorithm.


Short Communication on QUIST: A Quick Clustering Algorithm

arXiv.org Machine Learning

Then, it starts splitting C into smaller sub-clusters, until either one of the following conditions holds, whichever happens first: 1. The instances within each sub-cluster, c, are too similar to be divided any further, or 2. The number of clusters hits an optional upper bound provided by the user. If such bound is not provided by the user, it is assume to be equal to the number of input instances, C Additionally, splitting a given cluster stops when it reaches a minimum cluster size per the user's choice. To decide whether or not instances within a cluster, c, are similar enough to stop splitting it, QUIST calculates the "spreadness" metric denoted by Ψ, such that the spreadness of c, denoted by Ψ


Finding Singular Features

arXiv.org Machine Learning

We present a method for finding high density, low-dimensional structures in noisy point clouds. These structures are sets with zero Lebesgue measure with respect to the $D$-dimensional ambient space and belong to a $d