Statistical Learning
Stochastic Gradient Methods with Compressed Communication for Decentralized Saddle Point Problems
Sharma, Chhavi, Narayanan, Vishnu, Balamurugan, P.
We develop two compression based stochastic gradient algorithms to solve a class of non-smooth strongly convex-strongly concave saddle-point problems in a decentralized setting (without a central server). Our first algorithm is a Restart-based Decentralized Proximal Stochastic Gradient method with Compression (C-RDPSG) for general stochastic settings. We provide rigorous theoretical guarantees of C-RDPSG with gradient computation complexity and communication complexity of order $\mathcal{O}( (1+\delta)^4 \frac{1}{L^2}{\kappa_f^2}\kappa_g^2 \frac{1}{\epsilon} )$, to achieve an $\epsilon$-accurate saddle-point solution, where $\delta$ denotes the compression factor, $\kappa_f$ and $\kappa_g$ denote respectively the condition numbers of objective function and communication graph, and $L$ denotes the smoothness parameter of the smooth part of the objective function. Next, we present a Decentralized Proximal Stochastic Variance Reduced Gradient algorithm with Compression (C-DPSVRG) for finite sum setting which exhibits gradient computation complexity and communication complexity of order $\mathcal{O} \left((1+\delta) \max \{\kappa_f^2, \sqrt{\delta}\kappa^2_f\kappa_g,\kappa_g \} \log\left(\frac{1}{\epsilon}\right) \right)$. Extensive numerical experiments show competitive performance of the proposed algorithms and provide support to the theoretical results obtained.
Ranking Loss and Sequestering Learning for Reducing Image Search Bias in Histopathology
Mazaheri, Pooria, Bidgoli, Azam Asilian, Rahnamayan, Shahryar, Tizhoosh, H. R.
Recently, deep learning has started to play an essential role in healthcare applications, including image search in digital pathology. Despite the recent progress in computer vision, significant issues remain for image searching in histopathology archives. A well-known problem is AI bias and lack of generalization. A more particular shortcoming of deep models is the ignorance toward search functionality. The former affects every model, the latter only search and matching. Due to the lack of ranking-based learning, researchers must train models based on the classification error and then use the resultant embedding for image search purposes. Moreover, deep models appear to be prone to internal bias even if using a large image repository of various hospitals. This paper proposes two novel ideas to improve image search performance. First, we use a ranking loss function to guide feature extraction toward the matching-oriented nature of the search. By forcing the model to learn the ranking of matched outputs, the representation learning is customized toward image search instead of learning a class label. Second, we introduce the concept of sequestering learning to enhance the generalization of feature extraction. By excluding the images of the input hospital from the matched outputs, i.e., sequestering the input domain, the institutional bias is reduced. The proposed ideas are implemented and validated through the largest public dataset of whole slide images. The experiments demonstrate superior results compare to the-state-of-art.
Federated and distributed learning applications for electronic health records and structured medical data: A scoping review
Li, Siqi, Liu, Pinyan, Nascimento, Gustavo G., Wang, Xinru, Leite, Fabio Renato Manzolli, Chakraborty, Bibhas, Hong, Chuan, Ning, Yilin, Xie, Feng, Teo, Zhen Ling, Ting, Daniel Shu Wei, Haddadi, Hamed, Ong, Marcus Eng Hock, Peres, Marco Aurรฉlio, Liu, Nan
Federated learning (FL) has gained popularity in clinical research in recent years to facilitate privacy-preserving collaboration. Structured data, one of the most prevalent forms of clinical data, has experienced significant growth in volume concurrently, notably with the widespread adoption of electronic health records in clinical practice. This review examines FL applications on structured medical data, identifies contemporary limitations and discusses potential innovations. We searched five databases, SCOPUS, MEDLINE, Web of Science, Embase, and CINAHL, to identify articles that applied FL to structured medical data and reported results following the PRISMA guidelines. Each selected publication was evaluated from three primary perspectives, including data quality, modeling strategies, and FL frameworks. Out of the 1160 papers screened, 34 met the inclusion criteria, with each article consisting of one or more studies that used FL to handle structured clinical/medical data. Of these, 24 utilized data acquired from electronic health records, with clinical predictions and association studies being the most common clinical research tasks that FL was applied to. Only one article exclusively explored the vertical FL setting, while the remaining 33 explored the horizontal FL setting, with only 14 discussing comparisons between single-site (local) and FL (global) analysis. The existing FL applications on structured medical data lack sufficient evaluations of clinically meaningful benefits, particularly when compared to single-site analyses. Therefore, it is crucial for future FL applications to prioritize clinical motivations and develop designs and methodologies that can effectively support and aid clinical practice and research.
Multivariate regression modeling in integrative analysis via sparse regularization
Kawano, Shuichi, Fukushima, Toshikazu, Nakagawa, Junichi, Oshiki, Mamoru
The multivariate regression model basically offers the analysis of a single dataset with multiple responses. However, such a single-dataset analysis often leads to unsatisfactory results. Integrative analysis is an effective method to pool useful information from multiple independent datasets and provides better performance than single-dataset analysis. In this study, we propose a multivariate regression modeling in integrative analysis. The integration is achieved by sparse estimation that performs variable and group selection. Based on the idea of alternating direction method of multipliers, we develop its computational algorithm that enjoys the convergence property. The performance of the proposed method is demonstrated through Monte Carlo simulation and analyzing wastewater treatment data with microbe measurements.
Bayesian inference on Brain-Computer Interface using the GLASS Model
Zhao, Bangyao, Huggins, Jane E., Kang, Jian
The brain-computer interface (BCI) enables individuals with severe physical impairments to communicate with the world. BCIs offer computational neuroscience opportunities and challenges in converting real-time brain activities to computer commands and are typically framed as a classification problem. This article focuses on the P300 BCI that uses the event-related potential (ERP) BCI design, where the primary challenge is classifying target/non-target stimuli. We develop a novel Gaussian latent group model with sparse time-varying effects (GLASS) for making Bayesian inferences on the P300 BCI. GLASS adopts a multinomial regression framework that directly addresses the dataset imbalance in BCI applications. The prior specifications facilitate i) feature selection and noise reduction using soft-thresholding, ii) smoothing of the time-varying effects using global shrinkage, and iii) clustering of latent groups to alleviate high spatial correlations of EEG data. We develop an efficient gradient-based variational inference (GBVI) algorithm for posterior computation and provide a user-friendly Python module available at https://github.com/BangyaoZhao/GLASS. The application of GLASS identifies important EEG channels (PO8, Oz, PO7, Pz, C3) that align with existing literature. GLASS further reveals a group effect from channels in the parieto-occipital region (PO8, Oz, PO7), which is validated in cross-participant analysis.
Detection and Estimation of Structural Breaks in High-Dimensional Functional Time Series
Li, Degui, Li, Runze, Shang, Han Lin
Modelling functional time series, time series of random functions defined within a finite interval, has became one of the main frontiers of developments in time series models. Various functional linear and nonlinear time series models have been proposed and extensively studied in the past two decades (e.g., Bosq, 2000; Hรถrmann and Kokoszka, 2010; Horvรกth and Kokoszka, 2012; Hรถrmann, Horvรกth and Reeder, 2013; Li, Robinson and Shang, 2020). These models together with relevant methodologies have been applied to various fields such as biology, demography, economics, environmental science and finance. However, the model frameworks and methodologies developed in the aforementioned literature heavily rely on the stationarity assumption, which is often rejected when testing the functional time series data in practice. For example, Horvรกth, Kokoszka and Rice (2014) find evidence of nonstationarity for intraday price curves of some stocks collected in the US market; Aue, Rice and Sรถnmez (2018) reject the null hypothesis of stationarity for the temperature curves collected in Australia; and Li, Robinson and Shang (2023) reveal evidence of nonstationary feature for the functional time series constructed from the age-and sex-specific life-table death counts. It thus becomes imperative to test whether the collected functional time series are stationary. The primary interest of this paper is to test whether there exist structural breaks in the mean function over time and subsequently estimate locations of breaks if they do exist. There have been increasing interests on detecting and estimating structural breaks in functional time series. Broadly speaking, there are two types of detection techniques.
The Image of the M87 Black Hole Reconstructed with PRIMO - IOPscience
The exceptional resolution achieved by the EHT is made possible by an array of telescopes spanning the Earth and operating as a very long baseline interferometer (VLBI; Event Horizon Telescope Collaboration et al. 2019b, 2019c). Despite this global reach, the sparse interferometric coverage of the EHT array (especially during the 2017 observations that have been used for all of the publications to date) makes the already complex problem of interferometric image reconstruction particularly challenging. In such situations, special care is needed to assess the impact of imaging algorithms and sparse interferometric data on the final set of images that can be reconstructed from it. A cornerstone of the EHT data analysis strategy was the use of several independent analysis methods, each with different priorities, assumptions, and choices, to ensure that the EHT results were robust to these differences. The use of several general-purpose imaging algorithms, for example, was motivated by a desire to reconstruct an image that was consistent with the EHT data while remaining model-agnostic.
Use of machine learning algorithms to predict life-threatening ventricular arrhythmia in sepsis
Six ML algorithms including CatBoost, LightGBM, and XGBoost were employed to perform the model fitting. The least absolute shrinkage and selection operator (LASSO) regression was used to identify key features. Methods of model evaluation involved in this study included area under the receiver operating characteristic curve (AUROC), for model discrimination, calibration curve, and Brier score, for model calibration. A total of 27139 patients with sepsis were identified in this study, 1136 (4.2%) suffered from LTVA during hospitalization. We screened out 10 key features from the initial 54 variables via LASSO regression to improve the practicability of the model.
How to Use Loss Functions in TensorFlow
In this Vue tutorial, we learn about How to Use Loss Functions in TensorFlow. The loss metric is very important for neural networks. As all machine learning models are one optimization problem or another, the loss is the objective function to minimize. In neural networks, the optimization is done with gradient descent and backpropagation. But what are loss functions, and how are they affecting your neural networks?
The Evolution Of AI: Transforming The World One Algorithm At A Time
The journey of AI started in the 1950s with the pioneering work of Alan Turing, who proposed the Turing Test to determine if a machine could mimic human intelligence. In the 1960s, AI research gained momentum with the development of the first AI programming language, LISP, by John McCarthy. Early AI systems focused on symbolic reasoning and rule-based systems, which led to the development of expert systems in the 1970s and 1980s. The 1990s witnessed a shift in focus towards machine learning and data-driven approaches, driven by the increased availability of digital data and advancements in computing power. This period saw the rise of neural networks and the development of support vector machines, which allowed AI systems to learn from data, leading to better performance and adaptability.