Goto

Collaborating Authors

 Regression


Enhancement of Healthcare Data Performance Metrics using Neural Network Machine Learning Algorithms

arXiv.org Artificial Intelligence

Patients are often encouraged to make use of wearable devices for remote collection and monitoring of health data. This adoption of wearables results in a significant increase in the volume of data collected and transmitted. The battery life of the devices is then quickly diminished due to the high processing requirements of the devices. Given the importance attached to medical data, it is imperative that all transmitted data adhere to strict integrity and availability requirements. Reducing the volume of healthcare data for network transmission may improve sensor battery life without compromising accuracy. There is a trade-off between efficiency and accuracy which can be controlled by adjusting the sampling and transmission rates. This paper demonstrates that machine learning can be used to analyse complex health data metrics such as the accuracy and efficiency of data transmission to overcome the trade-off problem. The study uses time series nonlinear autoregressive neural network algorithms to enhance both data metrics by taking fewer samples to transmit. The algorithms were tested with a standard heart rate dataset to compare their accuracy and efficiency. The result showed that the Levenbery-Marquardt algorithm was the best performer with an efficiency of 3.33 and accuracy of 79.17%, which is similar to other algorithms accuracy but demonstrates improved efficiency. This proves that machine learning can improve without sacrificing a metric over the other compared to the existing methods with high efficiency.


Six online courses to learn regression in 2022

#artificialintelligence

Regression analysis is a useful mechanism for estimating the relationship between a dependent variable and one or more independent variables. It is widely used in forecasting and has become an important machine learning tool. It becomes crucial for someone starting in machine learning to understand how regression analysis works. Let us look at a few resources available online to get started with regression analysis. MachineHack, a popular platform for data scientists and AI practitioners provides courses on regression in the form of bootcamps. Bootcamps are pocket courses for all who aspire to become data scientists, data engineers and machine learning developers.


Machine Learning for Multi-Output Regression: When should a holistic multivariate approach be preferred over separate univariate ones?

arXiv.org Machine Learning

The hope of such multivariate analyses is, that the consideration of possible dependencies between the outcomes may lead to procedures with better power (in case of inference) or accuracy (in case of prediction) compared to separate univariate analyses. While the need for the development and use of valid and distributional robust or nonparametric multivariate methods has been recognized and addressed in inferential statistic (Dobler et al., 2020; Friedrich et al., 2019; Konietschke et al., 2015; Smaga, 2017; Vallejo and Ato, 2012; Zimmermann et al., 2020), there do not exist exhausting studies that exploit the potential of multivariate regression methods for prediction. Focussing on tree-based ensemble methods as the Random Forest, it is the aim of this manuscript to close this gap. In particular, we want to answer our research-motivating question: When should a holistic multivariate regression approach be preferred over separate univariate predictions? Corresponding Author Email address: lena.schmid@tu-dortmund.de (Lena Schmid)


A Kernel-Expanded Stochastic Neural Network

arXiv.org Machine Learning

The deep neural network suffers from many fundamental issues in machine learning. For example, it often gets trapped into a local minimum in training, and its prediction uncertainty is hard to be assessed. To address these issues, we propose the so-called kernel-expanded stochastic neural network (K-StoNet) model, which incorporates support vector regression (SVR) as the first hidden layer and reformulates the neural network as a latent variable model. The former maps the input vector into an infinite dimensional feature space via a radial basis function (RBF) kernel, ensuring absence of local minima on its training loss surface. The latter breaks the high-dimensional nonconvex neural network training problem into a series of low-dimensional convex optimization problems, and enables its prediction uncertainty easily assessed. The K-StoNet can be easily trained using the imputation-regularized optimization (IRO) algorithm. Compared to traditional deep neural networks, K-StoNet possesses a theoretical guarantee to asymptotically converge to the global optimum and enables the prediction uncertainty easily assessed. The performances of the new model in training, prediction and uncertainty quantification are illustrated by simulated and real data examples.


Synthesising Electronic Health Records: Cystic Fibrosis Patient Group

arXiv.org Artificial Intelligence

Class imbalance can often degrade predictive performance of supervised learning algorithms. Balanced classes can be obtained by oversampling exact copies, with noise, or interpolation between nearest neighbours (as in traditional SMOTE methods). Oversampling tabular data using augmentation, as is typical in computer vision tasks, can be achieved with deep generative models. Deep generative models are effective data synthesisers due to their ability to capture complex underlying distributions. Synthetic data in healthcare can enhance interoperability between healthcare providers by ensuring patient privacy. Equipped with large synthetic datasets which do well to represent small patient groups, machine learning in healthcare can address the current challenges of bias and generalisability. This paper evaluates synthetic data generators ability to synthesise patient electronic health records. We test the utility of synthetic data for patient outcome classification, observing increased predictive performance when augmenting imbalanced datasets with synthetic data.


Linear Regression in Python for Data Scientists

#artificialintelligence

Linear Regression is a statistical method used for modelling the dependence or relationship between two or more quantities. The aim of this is to be able to either better understand the existing relationships or to be able to predict the behaviour at points for which we currently don't have data. By using the method of linear regression (also called least squares fitting), we can calculate the values for the two parameters and plot the line of best fit to achieve our aims of better understanding the relationship or finding the estimated values of unknown points. For this, we have to be able to calculate the slope (m) and intercept (c) to give us the line of best fit for the data. This is made simple however by libraries that have already been implemented such as Scikit-Learn and Statsmodels Api that have linear regression functionality built in.


A robust kernel machine regression towards biomarker selection in multi-omics datasets of osteoporosis for drug discovery

arXiv.org Machine Learning

Many statistical machine approaches could ultimately highlight novel features of the etiology of complex diseases by analyzing multi-omics data. However, they are sensitive to some deviations in distribution when the observed samples are potentially contaminated with adversarial corrupted outliers (e.g., a fictional data distribution). Likewise, statistical advances lag in supporting comprehensive data-driven analyses of complex multi-omics data integration. We propose a novel non-linear M-estimator-based approach, "robust kernel machine regression (RobKMR)," to improve the robustness of statistical machine regression and the diversity of fictional data to examine the higher-order composite effect of multi-omics datasets. We address a robust kernel-centered Gram matrix to estimate the model parameters accurately. We also propose a robust score test to assess the marginal and joint Hadamard product of features from multi-omics data. We apply our proposed approach to a multi-omics dataset of osteoporosis (OP) from Caucasian females. Experiments demonstrate that the proposed approach effectively identifies the inter-related risk factors of OP. With solid evidence (p-value = 0.00001), biological validations, network-based analysis, causal inference, and drug repurposing, the selected three triplets ((DKK1, SMTN, DRGX), (MTND5, FASTKD2, CSMD3), (MTND5, COG3, CSMD3)) are significant biomarkers and directly relate to BMD. Overall, the top three selected genes (DKK1, MTND5, FASTKD2) and one gene (SIDT1 at p-value= 0.001) significantly bond with four drugs- Tacrolimus, Ibandronate, Alendronate, and Bazedoxifene out of 30 candidates for drug repurposing in OP. Further, the proposed approach can be applied to any disease model where multi-omics datasets are available.


Active Learning-Based Multistage Sequential Decision-Making Model with Application on Common Bile Duct Stone Evaluation

arXiv.org Machine Learning

Multistage sequential decision-making scenarios are commonly seen in the healthcare diagnosis process. In this paper, an active learning-based method is developed to actively collect only the necessary patient data in a sequential manner. There are two novelties in the proposed method. First, unlike the existing ordinal logistic regression model which only models a single stage, we estimate the parameters for all stages together. Second, it is assumed that the coefficients for common features in different stages are kept consistent. The effectiveness of the proposed method is validated in both a simulation study and a real case study. Compared with the baseline method where the data is modeled individually and independently, the proposed method improves the estimation efficiency by 62\%-1838\%. For both simulation and testing cohorts, the proposed method is more effective, stable, interpretable, and computationally efficient on parameter estimation. The proposed method can be easily extended to a variety of scenarios where decision-making can be done sequentially with only necessary information.


Leveraging Intrinsic Gradient Information for Further Training of Differentiable Machine Learning Models

arXiv.org Machine Learning

This work presents methods demonstrating that when the derivatives of target variables (outputs) with respect to inputs can be extracted - We introduce a novel metric that can be utilised in a from processes of interest, e.g., neural networks hyper-parameter optimisation pipeline that provides an (NN) based surrogate models, they can be leveraged indicator of an upper bound to NN model complexity to further improve the accuracy of differentiable - We propose an alternative regularisation method for linear ML models. This paper generalises the idea regression problems (using ridge regression as an and provides practical methodologies that can be example) that outperforms conventional regularisation used to leverage gradient information (GI) across over varying training sample sizes by utilising GI a variety of applications including: (1) Improving the performance of generative adversarial networks In the rest of this paper, Section 2 formulates the GI idea (GANs); (2) efficiently tuning NN model under a supervised learning setting. The proposed GI assisted complexity; (3) regularising linear regressions. Numerical methodologies are presented between Section 3 to 5, and followed results show that GI can effective enhance by a conclusion in Section 6. ML models with existing datasets, demonstrating its value for a variety of applications.


Generalized Shape Metrics on Neural Representations

arXiv.org Machine Learning

Understanding the operation of biological and artificial networks remains a difficult and important challenge. To identify general principles, researchers are increasingly interested in surveying large collections of networks that are trained on, or biologically adapted to, similar tasks. A standardized set of analysis tools is now needed to identify how network-level covariates -- such as architecture, anatomical brain region, and model organism -- impact neural representations (hidden layer activations). Here, we provide a rigorous foundation for these analyses by defining a broad family of metric spaces that quantify representational dissimilarity. Using this framework we modify existing representational similarity measures based on canonical correlation analysis to satisfy the triangle inequality, formulate a novel metric that respects the inductive biases in convolutional layers, and identify approximate Euclidean embeddings that enable network representations to be incorporated into essentially any off-the-shelf machine learning method. We demonstrate these methods on large-scale datasets from biology (Allen Institute Brain Observatory) and deep learning (NAS-Bench-101). In doing so, we identify relationships between neural representations that are interpretable in terms of anatomical features and model performance.