Goto

Collaborating Authors

 Genre


Logarithmic Time One-Against-Some

arXiv.org Machine Learning

We create a new online reduction of multiclass classification to binary classification for which training and prediction time scale logarithmically with the number of classes. Compared to previous approaches, we obtain substantially better statistical performance for two reasons: First, we prove a tighter and more complete boosting theorem, and second we translate the results more directly into an algorithm. We show that several simple techniques give rise to an algorithm that can compete with one-against-all in both space and predictive power while offering exponential improvements in speed when the number of classes is large.


Auditing Black-box Models for Indirect Influence

arXiv.org Machine Learning

Data-trained predictive models see widespread use, but for the most part they are used as black boxes which output a prediction or score. It is therefore hard to acquire a deeper understanding of model behavior, and in particular how different features influence the model prediction. This is important when interpreting the behavior of complex models, or asserting that certain problematic attributes (like race or gender) are not unduly influencing decisions. In this paper, we present a technique for auditing black-box models, which lets us study the extent to which existing models take advantage of particular features in the dataset, without knowing how the models work. Our work focuses on the problem of indirect influence: how some features might indirectly influence outcomes via other, related features. As a result, we can find attribute influences even in cases where, upon further direct examination of the model, the attribute is not referred to by the model at all. Our approach does not require the black-box model to be retrained. This is important if (for example) the model is only accessible via an API, and contrasts our work with other methods that investigate feature influence like feature selection. We present experimental evidence for the effectiveness of our procedure using a variety of publicly available datasets and models. We also validate our procedure using techniques from interpretable learning and feature selection, as well as against other black-box auditing procedures.


Classifiers for centrality determination in proton-nucleus and nucleus-nucleus collisions

arXiv.org Machine Learning

Centrality, as a geometrical property of the collision, is crucial for the physical interpretation of nucleus-nucleus and proton-nucleus experimental data. However, it cannot be directly accessed in event-by-event data analysis. Common methods for centrality estimation in AA and p-A collisions usually rely on a single detector (either on the signal in zero-degree calorimeters or on the multiplicity in some semi-central rapidity range). In the present work, we made an attempt to develop an approach for centrality determination that is based on machine-learning techniques and utilizes information from several detector subsystems simultaneously. Different event classifiers are suggested and evaluated for their selectivity power in terms of the number of nucleons-participants and the impact parameter of the collision. Finer centrality resolution may allow to reduce impact from so-called volume fluctuations on physical observables being studied in heavy-ion experiments like ALICE at the LHC and fixed target experiment NA61/SHINE on SPS. Machine-learning (ML) techniques have been used in High-Energy Physics (HEP) so far in a limited number of ways.


Stability selection for component-wise gradient boosting in multiple dimensions

arXiv.org Machine Learning

Noname manuscript No. (will be inserted by the editor) Abstract We present a new algorithm for boosting generalized additive models for location, scale and shape (GAMLSS) that allows to incorporate stability selection, an increasingly popular way to obtain stable sets of covariates while controlling the per-family error rate (PFER). The model is fitted repeatedly to subsampled data and variables with high selection frequencies are extracted. To apply stability selection to boosted GAMLSS, we develop a new "noncyclical" fitting algorithm that incorporates an additional selection step of the best-fitting distribution parameter in each iteration. This new algorithms has the additional advantage that optimizing the tuning parameters of boosting is reduced from a multidimensional to a one-dimensional problem with vastly decreased complexity. The performance of the novel algorithm is evaluated in an extensive simulation study. We apply this new algorithm to a study to estimate abundance of common eider in Massachusetts, USA, featuring excess zeros, overdispersion, non-linearity and spatiotemporal structures. Stability selection is used to obtain a sparse set of stable predictors. Keywords boosting ยท additive models ยท GAMLSS ยท gamboostLSS ยท Stability selection 1 Introduction In view of the growing size and complexity of modern databases, statistical modeling is increasingly faced with heteroscedasticity issues and a large number of available modeling options. In ecology, for example, it is often observed that outcome variables do not only show differences in mean conditions but also tend to be highly variable across different geographical features or states of a combination of covariates (e.g., [33]). In addition, ecological databases typically contain large numbers of correlated predictor variables that need to be carefully chosen for possible incorporation in a statistical regression model [1,8,31]. A convenient approach to address both heteroscedasticity and variable selection in statistical regression models is the combination of GAMLSS modeling with gradient boosting algorithms. GAMLSS, which refer to "generalized additive models for location, scale and shape" [34], are a modeling technique that relates not only the mean but all parameters of the outcome distribution to the available covariates.


Artificial Intelligence Vs Humans: AI More Likely To Enhance Life Rather Than Harm It

International Business Times

When it comes to artificial intelligence, you don't have to worry about killer robots taking over the world. In fact, AI promises to change our lives in ways we are beginning to experience, according to a survey produced by Stanford University. With a project called One Hundred Study on Artificial Intelligence (A100), Stanford is taking the long view of AI. The study, written by a panel of AI experts from various fields like healthcare, continue to release reports examining how AI will change different aspects of daily life. In the first report, Artificial Intelligence and Life in 2030, takes a look into the effects AI advancements will have on a North American city in a decade from now.


Accuracy of a Deep Learning Algorithm for Detection of Diabetic Retinopathy

#artificialintelligence

Question How does the performance of an automated deep learning algorithm compare with manual grading by ophthalmologists for identifying diabetic retinopathy in retinal fundus photographs? Finding In 2 validation sets of 9963 images and 1748 images, at the operating point selected for high specificity, the algorithm had 90.3% and 87.0% sensitivity and 98.1% and 98.5% specificity for detecting referable diabetic retinopathy, defined as moderate or worse diabetic retinopathy or referable macular edema by the majority decision of a panel of at least 7 US board-certified ophthalmologists. At the operating point selected for high sensitivity, the algorithm had 97.5% and 96.1% sensitivity and 93.4% and 93.9% specificity in the 2 validation sets. Meaning Deep learning algorithms had high sensitivity and specificity for detecting diabetic retinopathy and macular edema in retinal fundus photographs. Importance Deep learning is a family of computational methods that allow an algorithm to program itself by learning from a large set of examples that demonstrate the desired behavior, removing the need to specify rules explicitly. Application of these methods to medical imaging requires further assessment and validation. Objective To apply deep learning to create an algorithm for automated detection of diabetic retinopathy and diabetic macular edema in retinal fundus photographs. Design and Setting A specific type of neural network optimized for image classification called a deep convolutional neural network was trained using a retrospective development data set of 128 175 retinal images, which were graded 3 to 7 times for diabetic retinopathy, diabetic macular edema, and image gradability by a panel of 54 US licensed ophthalmologists and ophthalmology senior residents between May and December 2015.


Robots and the Future of Jobs: The Economic Impact of Artificial Intelligence

#artificialintelligence

I want to make one point, that this is on the record. But we're going to have a great time discussing "Robots and the Future of Jobs: The Economic Impact of Artificial Intelligence." So I'll start with simple introductions, and then we'll lay out some definitions about the kinds of terms that will be involved in this conversation. So my name is John Paul Farmer. Very happy to be here with three experts on the topic. Next to me is Dr. James Manyika, who is a recovering roboticist. And his day job is at McKinsey, at the McKinsey Global Institute, where he's been focusing on the future of jobs and the future of work in this new era. In the middle, we have Dr. Daniela Rus. Dr. Rus is a professor and roboticist at MIT, and she is also the director of the Computer Science and Artificial Intelligence Lab there. And at the end, we have Edwin van Bommel. Edwin is formerly of McKinsey, but now he's the chief cognitive officer at IPsoft. So, with that, let me lay out some definitions that are going to be important, I think, to following this conversation. You may have read in Foreign Affairs and elsewhere about this fourth industrial revolution, the changes that are happening in our society today and many more that will be coming down the pike. So as we--as we talk about these things, one, we should all be on the same page in terms of what artificial intelligence is. What do we mean when we say AI? And the definition that many accept is it's the development of computer systems able to perform tasks that normally require human intelligence, such as visual perception, speech recognition, decision-making, and even translation between languages. AI is sometimes humorously referred to as whatever computers can't do today. Machine learning is another term you're going to hear a lot, sometimes thought of as a rebranding of AI, of artificial intelligence. But there's one key difference, which is that it takes a much more probabilistic approach as opposed to deterministic. So it looks at not just yes or no; it looks at a 30 percent chance of X, a 10 percent chance of Y, and so on. Big data, a term that I think we've all heard. Data is the raw material. Some people call it the new oil for this new era.


Intel Elevates Its Autonomous Car Efforts Into A New Business Unit

Forbes - Tech

Nintendo Reports Second Quarter Losses But 3DS Sales Are Up Thanks To'Pokmon GO' Intel is reorganizing to better position itself for the next big thing in computing: self-driving cars. The chipmaking giant is taking its autonomous car efforts out the Internet of Things business group and creating a new business unit focused exclusively on the new market, called the Automated Driving Group. Doug Davis, the current head of Intel's Internet of Things division, will be heading up the new unit. In August, Davis had announced he would be retiring from Intel soon, but it looks like he changed his mind. "Throughout his career, Doug has consistently been on the leading side of disruption โ€“ standing up amazing new technologies that redefine how we experience work and life," said Intel president Murthy Renduchintala in a blog post.


Why Implement Machine Learning Algorithms From Scratch?

#artificialintelligence

Let us narrow down the phrase "implementing from scratch" a bit further in context of the 6 points I mentioned above. When we talk about "implementing from scratch," we need to narrow down the scope to make this question really tangible. Let's talk about a particular algorithm, simple logistic regression, to address the different points using concrete examples. I'd claim that logistic regression has been implemented more than thousand times. One reason why we'd still want to implement logistic regression from scratch could be that we don't have the impression that we fully understand how it works; we read a bunch of papers, and kind of understood the core concept though.


VA to employ artificial intelligence, precision medicine for veterans

#artificialintelligence

The Department of Veterans Affairs will work with Flow Health to build a medical knowledge graph to inform decision-making and train artificial intelligence to personalize care plans. The objective of the 5-year partnership, Flow Health executives said in a statement, is to understand the common elements that make certain people susceptible to particular diseases, to pinpoint effective treatments and identify possible side effects in order to inform care decisions. VA and Flow Health's work will entail integrating large volumes of data in the quest to discover relationships between genomes and phenotypes.The goal is to learn what every gene variant means, to identify disease risk, to make more precise diagnoses and to suggest individualized treatments. "Our mission is to advance healthcare by applying the latest artificial intelligence techniques to improve the detection, diagnosis, treatment and management of diseases," Flow Health CEO Alex Meshkin said in a statement. Flow Health is building a knowledge graph of medicine and genomics comprising more than 30 petabytes of longitudinal clinical data drawn from VA records on 22 million veterans spanning more than 20 years.