Genre
Machine Learning for Antimicrobial Resistance
Santerre, John W., Davis, James J., Xia, Fangfang, Stevens, Rick
Biological datasets amenable to applied machine learning are more available today than ever before, yet they lack adequate representation in the Data-for-Good community. Here we present a work in progress case study performing analysis on antimicrobial resistance (AMR) using standard ensemble machine learning techniques and note the successes and pitfalls such work entails. Broadly, applied machine learning (AML) techniques are well suited to AMR, with classification accuracies ranging from mid-90% to low- 80% depending on sample size. Additionally, these techniques prove successful at identifying gene regions known to be associated with the AMR phenotype. We believe that the extensive amount of biological data available, the plethora of problems presented, and the global impact of such work merits the consideration of the Data- for-Good community.
Sub-sampled Newton Methods with Non-uniform Sampling
Xu, Peng, Yang, Jiyan, Roosta-Khorasani, Farbod, Rรฉ, Christopher, Mahoney, Michael W.
We consider the problem of finding the minimizer of a convex function $F: \mathbb R^d \rightarrow \mathbb R$ of the form $F(w) := \sum_{i=1}^n f_i(w) + R(w)$ where a low-rank factorization of $\nabla^2 f_i(w)$ is readily available. We consider the regime where $n \gg d$. As second-order methods prove to be effective in finding the minimizer to a high-precision, in this work, we propose randomized Newton-type algorithms that exploit \textit{non-uniform} sub-sampling of $\{\nabla^2 f_i(w)\}_{i=1}^{n}$, as well as inexact updates, as means to reduce the computational complexity. Two non-uniform sampling distributions based on {\it block norm squares} and {\it block partial leverage scores} are considered in order to capture important terms among $\{\nabla^2 f_i(w)\}_{i=1}^{n}$. We show that at each iteration non-uniformly sampling at most $\mathcal O(d \log d)$ terms from $\{\nabla^2 f_i(w)\}_{i=1}^{n}$ is sufficient to achieve a linear-quadratic convergence rate in $w$ when a suitable initial point is provided. In addition, we show that our algorithms achieve a lower computational complexity and exhibit more robustness and better dependence on problem specific quantities, such as the condition number, compared to similar existing methods, especially the ones based on uniform sampling. Finally, we empirically demonstrate that our methods are at least twice as fast as Newton's methods with ridge logistic regression on several real datasets.
Building Ensembles of Adaptive Nested Dichotomies with Random-Pair Selection
Leathart, Tim, Pfahringer, Bernhard, Frank, Eibe
A system of nested dichotomies is a method of decomposing a multi-class problem into a collection of binary problems. Such a system recursively applies binary splits to divide the set of classes into two subsets, and trains a binary classifier for each split. Although ensembles of nested dichotomies with random structure have been shown to perform well in practice, using a more sophisticated class subset selection method can be used to improve classification accuracy. We investigate an approach to this problem called random-pair selection, and evaluate its effectiveness compared to other published methods of subset selection. We show that our method outperforms other methods in many cases when forming ensembles of nested dichotomies, and is at least on par in all other cases.
Sequential Dimensionality Reduction for Extracting Localized Features
Casalino, Gabriella, Gillis, Nicolas
Linear dimensionality reduction techniques are powerful tools for image analysis as they allow the identification of important features in a data set. In particular, nonnegative matrix factorization (NMF) has become very popular as it is able to extract sparse, localized and easily interpretable features by imposing an additive combination of nonnegative basis elements. Nonnegative matrix underapproximation (NMU) is a closely related technique that has the advantage to identify features sequentially. In this paper, we propose a variant of NMU that is particularly well suited for image analysis as it incorporates the spatial information, that is, it takes into account the fact that neighboring pixels are more likely to be contained in the same features, and favors the extraction of localized features by looking for sparse basis elements. We show that our new approach competes favorably with comparable state-of-the-art techniques on synthetic, facial and hyperspectral image data sets.
An Aggregate and Iterative Disaggregate Algorithm with Proven Optimality in Machine Learning
Park, Young Woong, Klabjan, Diego
In this paper, we propose a clustering-based iterative algorithm to solve certain optimization problems in machine learning when data size is large and thus it becomes impractical to use out-of-the-box algorithms. We rely on the principle of data aggregation and then subsequent disaggregations. While it is standard practice to aggregate the data and then calibrate the machine learning algorithm on aggregated data, we embed this into an iterative framework where initial aggregations are gradually disaggregated to the extent that even an optimal solution is obtainable. Early studies in data aggregation consider transportation problems [1, 10], where either demand or supply nodes are aggregated. Zipkin [31] studied data aggregation for linear programming (LP) and derived error bounds of the approximate solution.
Panasonic India to develop artificial intelligence tech for smartphones; plans strategic acquisitions ET Telecom
NEW DELHI: Panasonic India is scouting for companies to acquire in 6-9 months to develop artificial intelligence (AI) and Machine learning technologies, which it wants to integrate with its future smartphones to differentiate from rival vendors in the crowded yet fast growing market. The handset vendor has already set aside an initial corpus of 10 million for the development of this technology through a merger and acquisition or a joint venture. "The budget is in tune of 10 million to start with, and as we see progress on this front and things go in right direction, then there will be no constraint on the budget part. We can spend as high as possible. Some part of this budget has been generated from the India business, while some portion has been allocated from Japan," Pankaj Rana, head of mobility division, India, South Asia, Middle East and Africa at Panasonic, told ET. "Our team would be traveling to Silicon Valley soon. We will have new products ready with AI in 9-12 months. In the last three months, we have finalized what we will do and budgets have already been allocated from Panasonic Japan and Panasonic India. Now we have to find partner and start executive on timeline, while understanding the market," he said.
climate change big data ai
Climate deniers aside, there are very few people who are not concerned about the effects of climate change. According to a poll released by Monmouth University in January, nearly 70% of respondents said "the world's climate is undergoing a change leading to more extreme weather patterns and sea level rise",[1] and a more recent poll from Gallup reported 64% of "Americans are worried a great deal and/or a fair deal about global warming" [2]. What many people are unaware of is the fact climate scientists and business leaders are increasingly turning to Big Data and Artificial Intelligence (AI) to combat climate change. In many ways, this is inevitable as the sheer amount of data required to measure the effects of climate change requires the use of next-generation analytics. For example, the large data sets used to analyze climate are often prone to generating false positives and our understanding of climate change is still in its nascent stages.
Is AI The Worst Mistake In Human History?
One of the most intriguing public discussions to emerge over the past year is humanity's wrestling match with the threat and promise of artificial intelligence. AI has long lurked in our collective consciousness -- negatively so, if we're to take Hollywood movie plots as our guide -- but its recent andvery real advances are driving critical conversations about the future not only of our economy, but of humanity's very existence. In May 2014, the world received a wakeup call from famed physicist Stephen Hawking. Together with three respected AI researchers, the world's most renowned scientist warned that the commercially-driven creation of intelligent machines could be "potentially our worst mistake in history." Comparing the impact of AI on humanity to the arrival of "a superior alien species," Hawking and his co-authors found humanity's current state of preparedness deeply wanting.
New technologies are accelerating drug development, bringing hope to patients
In 2001, when Jamie was diagnosed with chronic myelogenous leukemia (CML), a cancer that starts inside the bone marrow, the disease had few effective cures. Fourteen years later, thanks to advances in cancer treatment, she is able to manage the disease and live a full life. Jamie is profiled in the PhRMA.org Yet many patients and their doctors wait for years before promising treatments become available. All too often, unforeseen side effects send researchers back to the drawing board, just when they thought they were close to bringing a new medication to market.
Robots could replace low-skilled migrant workers
Details of the fallout from the Brexit vote may take months to become clear, but there are concerns the UK pulling out of the European Union could lead to the loss of many low skilled migrant workers. But this apparent loss could be technology's gain, according to the findings from one think-tank. According to a new report from the Resolution Foundation, shortfalls in the human workforce could lead to a surge in robots to take their place. A new report from a think-tank says shortfalls in the workforce post-Brexit could lead to a surge in robots to take their place. According to findings from the Resolution Foundation, low-skilled jobs in agriculture and the food industry currently carried out by large numbers of EU workers could be automated.