Statistical Learning
AI Identifies Live Cancer Cells In Less Than 35 Minutes With 95% Accuracy
The ability to analyze single cells is one of the holy grails of precision medicine. Yuri Belotti, PhD, Doorgesh Sharma Jokhun, PhD, and Professor Chwee Teck (C.T.) Lim at National University of Singapore have developed a novel protocol for single-cell classification based on intracellular pH. Their paper entitled Machine learning based approach to pH imaging and classification of single cancer cells was published in APL Bioengineering. The pH in the human body varies between 4.7 and 8.0. Cancer growth, metastasis, and other diseases including Alzheimer's have been linked to deviations from normal intracellular acidity.
On the Optimality of the Oja's Algorithm for Online PCA
In this paper we analyze the behavior of the Oja's algorithm for online/streaming principal component subspace estimation. It is proved that with high probability it performs an efficient, gap-free, global convergence rate to approximate an principal component subspace for any sub-Gaussian distribution. Moreover, it is the first time to show that the convergence rate, namely the upper bound of the approximation, exactly matches the lower bound of an approximation obtained by the offline/classical PCA up to a constant factor.
AI-based Carcinoma Detection and Classification Using Histopathological Images: A Systematic Review
Prabhua, Swathi, Prasada, Keerthana, Robels-Kelly, Antonio, Lu, Xuequan
Histopathological image analysis is the gold standard to diagnose cancer. Carcinoma is a subtype of cancer that constitutes more than 80% of all cancer cases. Squamous cell carcinoma and adenocarcinoma are two major subtypes of carcinoma, diagnosed by microscopic study of biopsy slides. However, manual microscopic evaluation is a subjective and time-consuming process. Many researchers have reported methods to automate carcinoma detection and classification. The increasing use of artificial intelligence (AI) in the automation of carcinoma diagnosis also reveals a significant rise in the use of deep network models. In this systematic literature review, we present a comprehensive review of the state-of-the-art approaches reported in carcinoma diagnosis using histopathological images. Studies are selected from well-known databases with strict inclusion/exclusion criteria. We have categorized the articles and recapitulated their methods based on specific organs of carcinoma origin. Further, we have summarized pertinent literature on AI methods, highlighted critical challenges and limitations, and provided insights on future research direction in automated carcinoma diagnosis. Out of 101 articles selected, most of the studies experimented on private datasets with varied image sizes, obtaining accuracy between 63% and 100%. Overall, this review highlights the need for a generalized AI-based carcinoma diagnostic system. Additionally, it is desirable to have accountable approaches to extract microscopic features from images of multiple magnifications that should mimic pathologists' evaluations.
WATCH: Wasserstein Change Point Detection for High-Dimensional Time Series Data
Faber, Kamil, Corizzo, Roberto, Sniezynski, Bartlomiej, Baron, Michael, Japkowicz, Nathalie
Detecting relevant changes in dynamic time series data in a timely manner is crucially important for many data analysis tasks in real-world settings. Change point detection methods have the ability to discover changes in an unsupervised fashion, which represents a desirable property in the analysis of unbounded and unlabeled data streams. However, one limitation of most of the existing approaches is represented by their limited ability to handle multivariate and high-dimensional data, which is frequently observed in modern applications such as traffic flow prediction, human activity recognition, and smart grids monitoring. In this paper, we attempt to fill this gap by proposing WATCH, a novel Wasserstein distance-based change point detection approach that models an initial distribution and monitors its behavior while processing new data points, providing accurate and robust detection of change points in dynamic high-dimensional data. An extensive experimental evaluation involving a large number of benchmark datasets shows that WATCH is capable of accurately identifying change points and outperforming state-of-the-art methods.
Inducing Structure in Reward Learning by Learning Features
Bobu, Andreea, Wiggert, Marius, Tomlin, Claire, Dragan, Anca D.
In doing so, however, these approaches sacrifice the sample efficiency and generalizability that a well-specified feature Whether it's semi-autonomous driving (Sadigh et al. 2016), set offers. While using an expressive function approximator recommender systems (Ziebart et al. 2008), or household to extract features and learn their reward combination at once robots working in close proximity with people (Jain et al. seems advantageous, many such functions can induce policies 2015), reward learning can greatly benefit autonomous agents that explain the demonstrations. Hence, to disambiguate to generate behaviors that adapt to new situations or human between all these candidate functions, the robot requires a preferences. Under this framework, the robot uses the person's very large amount of (laborious to collect) data, and this data input to learn a reward function that describes how they prefer needs to be diverse enough to identify the true reward. For the task to be performed. For instance, in the scenario in Fig. example, the human in the household robot setting in Figure 1 1, the human wants the robot to keep the cup away from the might want to demonstrate keeping the cup away from the laptop to prevent spilling liquid over it; she may communicate laptop, but from a single demonstration the robot could find this preference to the robot by providing a demonstration of many other explanations for the person's behavior: perhaps the task or even by directly intervening during the robot's task they always happened to keep the cup upright or they really execution to correct it.
Representation Learning on Heterostructures via Heterogeneous Anonymous Walks
Guo, Xuan, Jiao, Pengfei, Pan, Ting, Zhang, Wang, Jia, Mengyu, Shi, Danyang, Wang, Wenjun
Capturing structural similarity has been a hot topic in the field of network embedding recently due to its great help in understanding the node functions and behaviors. However, existing works have paid very much attention to learning structures on homogeneous networks while the related study on heterogeneous networks is still a void. In this paper, we try to take the first step for representation learning on heterostructures, which is very challenging due to their highly diverse combinations of node types and underlying structures. To effectively distinguish diverse heterostructures, we firstly propose a theoretically guaranteed technique called heterogeneous anonymous walk (HAW) and its variant coarse HAW (CHAW). Then, we devise the heterogeneous anonymous walk embedding (HAWE) and its variant coarse HAWE in a data-driven manner to circumvent using an extremely large number of possible walks and train embeddings by predicting occurring walks in the neighborhood of each node. Finally, we design and apply extensive and illustrative experiments on synthetic and real-world networks to build a benchmark on heterostructure learning and evaluate the effectiveness of our methods. The results demonstrate our methods achieve outstanding performance compared with both homogeneous and heterogeneous classic methods, and can be applied on large-scale networks.
Multiway Spherical Clustering via Degree-Corrected Tensor Block Models
We consider the problem of multiway clustering in the presence of unknown degree heterogeneity. Such data problems arise commonly in applications such as recommendation system, neuroimaging, community detection, and hypergraph partitions in social networks. The allowance of degree heterogeneity provides great flexibility in clustering models, but the extra complexity poses significant challenges in both statistics and computation. Here, we develop a degree-corrected tensor block model with estimation accuracy guarantees. We present the phase transition of clustering performance based on the notion of angle separability, and we characterize three signal-to-noise regimes corresponding to different statistical-computational behaviors. In particular, we demonstrate that an intrinsic statistical-to-computational gap emerges only for tensors of order three or greater. Further, we develop an efficient polynomial-time algorithm that provably achieves exact clustering under mild signal conditions. The efficacy of our procedure is demonstrated through two data applications, one on human brain connectome project, and another on Peru Legislation network dataset.
Learning Tensor Representations for Meta-Learning
Deng, Samuel, Guo, Yilin, Hsu, Daniel, Mandal, Debmalya
We introduce a tensor-based model of shared representation for meta-learning from a diverse set of tasks. Prior works on learning linear representations for meta-learning assume that there is a common shared representation across different tasks, and do not consider the additional task-specific observable side information. In this work, we model the meta-parameter through an order-$3$ tensor, which can adapt to the observed task features of the task. We propose two methods to estimate the underlying tensor. The first method solves a tensor regression problem and works under natural assumptions on the data generating process. The second method uses the method of moments under additional distributional assumptions and has an improved sample complexity in terms of the number of tasks. We also focus on the meta-test phase, and consider estimating task-specific parameters on a new task. Substituting the estimated tensor from the first step allows us estimating the task-specific parameters with very few samples of the new task, thereby showing the benefits of learning tensor representations for meta-learning. Finally, through simulation and several real-world datasets, we evaluate our methods and show that it improves over previous linear models of shared representations for meta-learning.
Online, Informative MCMC Thinning with Kernelized Stein Discrepancy
Hawkins, Cole, Koppel, Alec, Zhang, Zheng
A fundamental challenge in Bayesian inference is efficient representation of a target distribution. Many non-parametric approaches do so by sampling a large number of points using variants of Markov Chain Monte Carlo (MCMC). We propose an MCMC variant that retains only those posterior samples which exceed a KSD threshold, which we call KSD Thinning. We establish the convergence and complexity tradeoffs for several settings of KSD Thinning as a function of the KSD threshold parameter, sample size, and other problem parameters. Finally, we provide experimental comparisons against other online nonparametric Bayesian methods that generate low-complexity posterior representations, and observe superior consistency/complexity tradeoffs. Code is available at github.com/colehawkins/KSD-Thinning.
Socioeconomic disparities and COVID-19: the causal connections
Banerjee, Tannista, Paul, Ayan, Srikanth, Vishak, Strümke, Inga
The analysis of causation is a challenging task that can be approached in various ways. With the increasing use of machine learning based models in computational socioeconomics, explaining these models while taking causal connections into account is a necessity. In this work, we advocate the use of an explanatory framework from cooperative game theory augmented with $do$ calculus, namely causal Shapley values. Using causal Shapley values, we analyze socioeconomic disparities that have a causal link to the spread of COVID-19 in the USA. We study several phases of the disease spread to show how the causal connections change over time. We perform a causal analysis using random effects models and discuss the correspondence between the two methods to verify our results. We show the distinct advantages a non-linear machine learning models have over linear models when performing a multivariate analysis, especially since the machine learning models can map out non-linear correlations in the data. In addition, the causal Shapley values allow for including the causal structure in the variable importance computed for the machine learning model.