Goto

Collaborating Authors

 Genre


Otto Product Classification Winner's Interview: 2nd place, Alexander Guschin \_(?)_/

#artificialintelligence

The Otto Group Product Classification Challenge made Kaggle history as our most popular competition ever. Alexander Guschin finished in 2nd place ahead of 3,845 other data scientists. In this blog, Alexander shares his stacking centered approach and explains why you should never underestimate the nearest neighbours algorithm. I have some theoretical understanding of machine learning thanks to my base institute (Moscow Institute of Physics and Technology) and our professor Konstantin Vorontsov, one of the top Russian machine learning specialists. As for my acquaintance with practical problems, another great Russian data scientist who once was Top-1 on Kaggle, Alexander D'yakonov, used to teach a course on practical machine learning every autumn which gave me very good basis. Kagglers may know this course as PZAD.


Hitachi : June 27, 2016Hitachi Develops Technology to Automatically Create Effective Advice to Increase Worker Happiness Using Artificial Intelligence 4-Traders

#artificialintelligence

Tokyo, June 27, 2016 --- Hitachi, Ltd. (TSE:6501; 'Hitachi') today announced the development of technology using artificial intelligence that automatically creates effective advice for raising the happiness of workers based on the behavioral data of each individual on a daily basis and the commencement of an internal trial with 600 participants from sales & marketing. More precisely, a name tag type wearable sensor collects the massive amount of individual behavioral data which is then analyzed using Hitachi AI*1Technology/H (hereafter referred to as H), and used to create and deliver personalized advice on actions automatically, such as on communication in the workplace or time allocation that will contribute to raising individual happiness. This advice is delivered daily, and individual workers can check the daily advice on their smartphone or tablet, and choose to apply the advice in their daily activity. Hitachi will integrate the results from this trial into the solution which it will provide to corporate and other organizations globally, to support increased productivity through a more active organization resulting from the increased happiness of workers. In recent years, increasing'happiness' has become one of society's most important issues.


Understanding the impact of AI

#artificialintelligence

Coding will join this list in time, however, where it differs wildly from the afore mentioned examples is it is unlikely to be lovingly preserved for future generations to admire, fiddle with or better still, reactivate. Its essence will not be reified for one specific reason โ€“ it can't be touched and humans value tactility. We touch immediately, both inside and outside the womb. Today, we find ourselves at a pivotal moment in our existence and about to experience an exponential period of rapid technological growth the likes of which is quite probably beyond our comprehension and at a base level, will have serious implications for coding. We rather arrogantly think that because we have a good grasp of our own technological advancement so far, we can somehow predict the mass cultural and behavioural shift about to happen as we question our own skills in the world.


Study exposes major flaw in classic artificial intelligence test

#artificialintelligence

A serious problem in the Turing test for computer intelligence is exposed in a study published in the Journal of Experimental and Theoretical Artificial Intelligence. If a machine were to'take the Fifth Amendment' โ€“ that is, exercise the right to remain silent throughout the test โ€“ it could, potentially, pass the test and thus be regarded as a thinking entity, authors Kevin Warwick and Huma Shah of Coventry University argue. However, if this is the case, any silent entity could pass the test, even if it were clearly incapable of thought. The test, devised in 1950 by pioneering computer scientist Alan Turing, assesses a machine's ability to exhibit intelligent behaviour indistinguishable from that of a human. Also known as the'imitation game', it requires a human judge to converse with two hidden entities, a human and a machine, and then determine which is which.


A look at AirBnB demographics R-bloggers

#artificialintelligence

Once in a while I use AirBnB. There are a couple of features that I (intuitively) use to judge if an apartment is save to book; ratings, images of the flat and the user avatar. Apparently, these avatars play an important part in the overall service and usage of AirBnB. A recent study finds that "Attractive Airbnb hosts are more likely to get bookings, even with bad reviews". With the easy availability of image recognition services, even the everyday researcher can do a small analysis.


Optimal control for a robotic exploration, pick-up and delivery problem

arXiv.org Artificial Intelligence

Different versions of this problem have received is coping with uncertainties arising from limited a-considerable attention from several research communities, e.g., priori knowledge of the environment. Acquiring necessary as a "pursuit-evasion game" in game theory [13], [14], as information and achieving the overall goal are complementary a "cow-path problem" in computer science [15] or as a subtasks that require adapting the motion of a robot during "coverage problem" in control [16], [17], but its solution for mission execution, typically accompanied by minimizing a a general probability distribution or a general geometry of the performance criterion. In this work we address an Optimal region is, to a large extent, still an open question. Effective Control Problem (OCP) for a robot with fourth-order dynamics approaches for the related persistent monitoring problem based that has to find, collect and move a finite number of on estimation [18], linear programming [19] or parametric objects to a designated spot in minimum time. The objects optimization [20] have been also been proposed. OCPs with with a-priori known masses are located in a bounded twodimensional uncertainties have also been addressed by certainty equivalent space, where the robot is capable of localizing event-triggered [21], minimax [22] and sampling-based [23] itself using a state-of-the-art simultaneous localization and optimization schemes.


How to Evaluate the Quality of Unsupervised Anomaly Detection Algorithms?

arXiv.org Machine Learning

When sufficient labeled data are available, classical criteria based on Receiver Operating Characteristic (ROC) or Precision-Recall (PR) curves can be used to compare the performance of un-supervised anomaly detection algorithms. However , in many situations, few or no data are labeled. This calls for alternative criteria one can compute on non-labeled data. In this paper, two criteria that do not require labels are empirically shown to discriminate accurately (w.r.t. ROC or PR based criteria) between algorithms. These criteria are based on existing Excess-Mass (EM) and Mass-Volume (MV) curves, which generally cannot be well estimated in large dimension. A methodology based on feature sub-sampling and aggregating is also described and tested, extending the use of these criteria to high-dimensional datasets and solving major drawbacks inherent to standard EM and MV curves.


An optimal learning method for developing personalized treatment regimes

arXiv.org Machine Learning

A treatment regime is a function that maps individual patient information to a recommended treatment, hence explicitly incorporating the heterogeneity in need for treatment across individuals. Patient responses are dichotomous and can be predicted through an unknown relationship that depends on the patient information and the selected treatment. The goal is to find the treatments that lead to the best patient responses on average. Each experiment is expensive, forcing us to learn the most from each experiment. We adopt a Bayesian approach both to incorporate possible prior information and to update our treatment regime continuously as information accrues, with the potential to allow smaller yet more informative trials and for patients to receive better treatment. By formulating the problem as contextual bandits, we introduce a knowledge gradient policy to guide the treatment assignment by maximizing the expected value of information, for which an approximation method is used to overcome computational challenges. We provide a detailed study on how to make sequential medical decisions under uncertainty to reduce health care costs on a real world knee replacement dataset. We use clustering and LASSO to deal with the intrinsic sparsity in health datasets. We show experimentally that even though the problem is sparse, through careful selection of physicians (versus picking them at random), we can significantly improve the success rates.


One-Shot Session Recommendation Systems with Combinatorial Items

arXiv.org Machine Learning

In recent years, content recommendation systems in large websites (or \emph{content providers}) capture an increased focus. While the type of content varies, e.g.\ movies, articles, music, advertisements, etc., the high level problem remains the same. Based on knowledge obtained so far on the user, recommend the most desired content. In this paper we present a method to handle the well known user-cold-start problem in recommendation systems. In this scenario, a recommendation system encounters a new user and the objective is to present items as relevant as possible with the hope of keeping the user's session as long as possible. We formulate an optimization problem aimed to maximize the length of this initial session, as this is believed to be the key to have the user come back and perhaps register to the system. In particular, our model captures the fact that a single round with low quality recommendation is likely to terminate the session. In such a case, we do not proceed to the next round as the user leaves the system, possibly never to seen again. We denote this phenomenon a \emph{One-Shot Session}. Our optimization problem is formulated as an MDP where the action space is of a combinatorial nature as we recommend in each round, multiple items. This huge action space presents a computational challenge making the straightforward solution intractable. We analyze the structure of the MDP to prove monotone and submodular like properties that allow a computationally efficient solution via a method denoted by \emph{Greedy Value Iteration} (G-VI).


Efficient Estimation in the Tails of Gaussian Copulas

arXiv.org Machine Learning

We consider the question of efficient estimation in the tails of Gaussian copulas. Our special focus is estimating expectations over multi-dimensional constrained sets that have a small implied measure under the Gaussian copula. We propose three estimators, all of which rely on a simple idea: identify certain \emph{dominating} point(s) of the feasible set, and appropriately shift and scale an exponential distribution for subsequent use within an importance sampling measure. As we show, the efficiency of such estimators depends crucially on the local structure of the feasible set around the dominating points. The first of our proposed estimators $\estOpt$ is the "full-information" estimator that actively exploits such local structure to achieve bounded relative error in Gaussian settings. The second and third estimators $\estExp$, $\estLap$ are "partial-information" estimators, for use when complete information about the constraint set is not available, they do not exhibit bounded relative error but are shown to achieve polynomial efficiency. We provide sharp asymptotics for all three estimators. For the NORTA setting where no ready information about the dominating points or the feasible set structure is assumed, we construct a multinomial mixture of the partial-information estimator $\estLap$ resulting in a fourth estimator $\estNt$ with polynomial efficiency, and implementable through the ecoNORTA algorithm. Numerical results on various example problems are remarkable, and consistent with theory.