Goto

Collaborating Authors

 Statistical Learning


Steady Progress of AI Makes its Mark in the Clinic

#artificialintelligence

A 65-year-old man with a remote smoking history is automatically referred to your clinic after an artificial intelligence (AI) algorithm-enhanced routine CT screening exam detects a nodule with 90% certainty of malignancy. A 45-year-old woman with metastatic breast cancer arrives at your office and her AI-integrated electronic health record (EHR) algorithm alerts you that she has a 60% probability of a major cardiac event within 5 years if given an anthracycline. You hold an uninterrupted, computer screen–free discussion with your patient about their newly diagnosed pancreatic cancer; meanwhile, an AI-powered device processes the conversation and produces all documentation and billing information. Only a few years ago, these imagined scenarios may have seemed far-fetched, but no longer. Increasingly, and with exponential pace, AI algorithms are finding their way into the oncology clinic.


Top 7 beginner projects in Machine Learning

#artificialintelligence

Well, I hope this article has helped you understand models and their codes. You may use the link to search for a model and learn more about it.


Building a Random Forest Classifier to Predict Neural Spikes

#artificialintelligence

A step-by-step guide to building a Random Forest classifier in Python to predict subtypes of neural extracellular spikes using a real data-set recorded from Human brain organoids. Given the heterogeneity of neurons within the human brain itself, classification tools are commonly utilised to correlate electrical activity with different cell types and/or morphologies. This is a long-standing question in Neuroscience circles, and can be considerably variable between different species, pathologies, brain regions and layers. Fortunately, with the readily increasing computational power allowing improvements in machine-learning and deep-learning algorithms, Neuroscientists are provided with the tools to dive further into asking these important questions. However, as stated by Juavinett et al., for the most part programming skills are underrepresented in the community and new resources to teach them are crucial to solving the complexity of the human brain.


Multiclass Classification Using TensorFlow

#artificialintelligence

In the previous article, I discussed building a linear regression model using Tensorflow. In this article, I will try to solve a multiclass classification problem using Tensorflow. I have used the MNIST-digit recognizer dataset here. Please note that even though a Convolutional Neural Network might have worked better for this problem as this is an image recognition problem, but I have used a generic neural network as I wanted to showcase solving a classification problem using Neural Networks. The dataset consists of 784 pixel columns, where each row represents a 28 x 28 image flattened out into a row vector, and a label column, with the image labels given by the digits the image represent, from 0–9.


Implicit Regularization Towards Rank Minimization in ReLU Networks

arXiv.org Machine Learning

A central puzzle in the theory of deep learning is how neural networks generalize even when trained without any explicit regularization, and when there are far more learnable parameters than training examples. In such an underdetermined optimization problem, there are many global minima with zero training loss, and gradient descent seems to prefer solutions that generalize well (see Zhang et al. (2017)). Hence, it is believed that gradient descent induces an implicit regularization (or implicit bias) (Neyshabur et al., 2015, 2017), and characterizing this regularization/bias has been a subject of extensive research. Several works in recent years studied the relationship between the implicit regularization in linear neural networks and rank minimization. A main focus is on the matrix factorization problem, which corresponds to training a depth-2 linear neural network with multiple outputs w.r.t. the square loss, and is considered a well-studied test-bed for studying implicit regularization in deep learning.


Heterogeneous Federated Learning via Grouped Sequential-to-Parallel Training

arXiv.org Artificial Intelligence

Federated learning (FL) is a rapidly growing privacy-preserving collaborative machine learning paradigm. In practical FL applications, local data from each data silo reflect local usage patterns. Therefore, there exists heterogeneity of data distributions among data owners (a.k.a. FL clients). If not handled properly, this can lead to model performance degradation. This challenge has inspired the research field of heterogeneous federated learning, which currently remains open. In this paper, we propose a data heterogeneity-robust FL approach, FedGSP, to address this challenge by leveraging on a novel concept of dynamic Sequential-to-Parallel (STP) collaborative training. FedGSP assigns FL clients to homogeneous groups to minimize the overall distribution divergence among groups, and increases the degree of parallelism by reassigning more groups in each round. It is also incorporated with a novel Inter-Cluster Grouping (ICG) algorithm to assist in group assignment, which uses the centroid equivalence theorem to simplify the NP-hard grouping problem to make it solvable. Extensive experiments have been conducted on the non-i.i.d. FEMNIST dataset. The results show that FedGSP improves the accuracy by 3.7% on average compared with seven state-of-the-art approaches, and reduces the training time and communication overhead by more than 90%.


MVP-Net: Multiple View Pointwise Semantic Segmentation of Large-Scale Point Clouds

arXiv.org Artificial Intelligence

Semantic segmentation of 3D point cloud is an essential task for autonomous driving environment perception. The pipeline of most pointwise point cloud semantic segmentation methods includes points sampling, neighbor searching, feature aggregation, and classification. Neighbor searching method like K-nearest neighbors algorithm, KNN, has been widely applied. However, the complexity of KNN is always a bottleneck of efficiency. In this paper, we propose an end-to-end neural architecture, Multiple View Pointwise Net, MVP-Net, to efficiently and directly infer large-scale outdoor point cloud without KNN or any complex pre/postprocessing. Instead, assumption-based sorting and multi-rotation of point cloud methods are introduced to point feature aggregation and receptive field expanding. Numerical experiments show that the proposed MVP-Net is 11 times faster than the most efficient pointwise semantic segmentation method RandLA-Net and achieves the same accuracy on the large-scale benchmark SemanticKITTI dataset.


Homotopic Policy Mirror Descent: Policy Convergence, Implicit Regularization, and Improved Sample Complexity

arXiv.org Artificial Intelligence

We propose the homotopic policy mirror descent (HPMD) method for solving discounted, infinite horizon MDPs with finite state and action space, and study its policy convergence. We report three properties that seem to be new in the literature of policy gradient methods: (1) HPMD exhibits global linear convergence of the value optimality gap, and local superlinear convergence of the policy to the set of optimal policies with order $\gamma^{-2}$. The superlinear convergence of the policy takes effect after no more than $\mathcal{O}(\log(1/\Delta^*))$ number of iterations, where $\Delta^*$ is defined via a gap quantity associated with the optimal state-action value function; (2) HPMD also exhibits last-iterate convergence of the policy, with the limiting policy corresponding exactly to the optimal policy with the maximal entropy for every state. No regularization is added to the optimization objective and hence the second observation arises solely as an algorithmic property of the homotopic policy gradient method. (3) For the stochastic HPMD method, we further demonstrate a better than $\mathcal{O}(|\mathcal{S}| |\mathcal{A}| / \epsilon^2)$ sample complexity for small optimality gap $\epsilon$, when assuming a generative model for policy evaluation.


Provable Domain Generalization via Invariant-Feature Subspace Recovery

arXiv.org Machine Learning

Domain generalization asks for models trained on a set of training environments to perform well on unseen test environments. Recently, a series of algorithms such as Invariant Risk Minimization (IRM) has been proposed for domain generalization. However, Rosenfeld et al. (2021) shows that in a simple linear data model, even if non-convexity issues are ignored, IRM and its extensions cannot generalize to unseen environments with less than $d_s+1$ training environments, where $d_s$ is the dimension of the spurious-feature subspace. In this paper, we propose to achieve domain generalization with Invariant-feature Subspace Recovery (ISR). Our first algorithm, ISR-Mean, can identify the subspace spanned by invariant features from the first-order moments of the class-conditional distributions, and achieve provable domain generalization with $d_s+1$ training environments under the data model of Rosenfeld et al. (2021). Our second algorithm, ISR-Cov, further reduces the required number of training environments to $O(1)$ using the information of second-order moments. Notably, unlike IRM, our algorithms bypass non-convexity issues and enjoy global convergence guarantees. Empirically, our ISRs can obtain superior performance compared with IRM on synthetic benchmarks. In addition, on three real-world image and text datasets, we show that ISR-Mean can be used as a simple yet effective post-processing method to increase the worst-case accuracy of trained models against spurious correlations and group shifts.


Approximate Bayesian Computation Based on Maxima Weighted Isolation Kernel Mapping

arXiv.org Machine Learning

This paper addresses the problem of precisely estimating the parameters of a stochastic model corresponding to branching processes. A branching process is a stochastic process consisting of collections of random variables indexed by the natural numbers. Branching processes are often used to describe population models Jagers (1989) and Athreya and Ney (2012); for example, models in the population genetics showing the genetic drift Burden and Simon (2016) Chen et al. (2017). In contrast to statistical approaches, branching processes enable the study of the dynamics of cell evolution and, as a consistence, have become a popular approach to cancer cell evolution research West et al., 2016. However, particularly in the case of cancer cell evolution, as well as in branching processes in general, the ultimate extinction of a population often occurs Devroye (1998). It is for this reason that with the initial uniform distribution of parameters, branching processes models tend to yield unevenly distributed data consisting of sparse and dense regions. The stochastic nature of the data is an another obstacle in estimating the parameters of a branching processes model, especially in the case of cancer cell evolution Nagornov et al. (2021). Moreover, simulations, based on a model of cell mutations, population evolution, and tumor/cancer subpopulations, commonly lead to the emergence of many clones and rarely to the appearance of cancer cells.