Statistical Learning
Better Model Performance Through Algorithmic Diversity - DZone AI
The upcoming release on BigML keeps pushing the envelope for more people to gain access to and make a bigger impact with Machine Learning. This release features a brand new resource providing a novel way to combine models: BigML Fusions. In this post, we'll do a quick introduction to Fusions before we move on to the remainder of our series of 6 blog posts (including this one) to give you a detailed perspective of what's behind this new capability. Today's post explains the basic concepts that will be followed by an example use case. This will be followed by three more articles focused on how to use Fusions through the BigML Dashboard, API, and WhizzML in an automated fashion.
VFPred: A Fusion of Signal Processing and Machine Learning techniques in Detecting Ventricular Fibrillation from ECG Signals
Ibtehaz, Nabil, Rahman, M. Saifur, Rahman, M. Sohel
Ventricular Fibrillation (VF), one of the most dangerous arrhythmias, is responsible for sudden cardiac arrests. Thus, various algorithms have been developed to predict VF from Electrocardiogram (ECG), which is a binary classification problem. In the literature, we find a number of algorithms based on signal processing, where, after some robust mathematical operations the decision is given based on a predefined threshold over a single value. On the other hand, some machine learning based algorithms are also reported in the literature; however, these algorithms merely combine some parameters and make a prediction using those as features. Both the approaches have their perks and pitfalls; thus our motivation was to coalesce them to get the best out of the both worlds. Sohel Rahman) Preprint submitted to Pattern Recognition July 10, 2018 a Support Vector Machine for efficient classification. VFPred turns out to be a robust algorithm as it is able to successfully segregate the two classes with equal confidence (Sensitivity 99.99%, Specificity 98.40%) even from a short signal of 5 seconds long, whereas existing works though requires longer signals, flourishes in one but fails in the other. Keywords: Electrocardiogram(ECG), Empirical Mode Decomposition, Heart Arrhythmia, Support Vector Machine, Ventricular Fibrillation(VF). 1. Introduction Ventricular Fibrillation (VF) is a type of cardiac arrhythmia which occurs when the heart quivers instead of pumping due to disturbance in electrical activity in the ventricles [1]. This arrhythmia may result in a cardiac arrest leaving the patient unconscious without any pulse. Ventricular Fibrillation is found initially in about 10% of people in cardiac arrest [2] and sudden cardiac arrest is responsible for approximately 6 million deaths in Europe and in the United States [3]. Therefore, fast and accurate detection of Ventricular Fibrillation can save a lot of lives.
Approximate Leave-One-Out for Fast Parameter Tuning in High Dimensions
Wang, Shuaiwen, Zhou, Wenda, Lu, Haihao, Maleki, Arian, Mirrokni, Vahab
Consider the following class of learning schemes: $$\hat{\boldsymbol{\beta}} := \arg\min_{\boldsymbol{\beta}}\;\sum_{j=1}^n \ell(\boldsymbol{x}_j^\top\boldsymbol{\beta}; y_j) + \lambda R(\boldsymbol{\beta}),\qquad\qquad (1) $$ where $\boldsymbol{x}_i \in \mathbb{R}^p$ and $y_i \in \mathbb{R}$ denote the $i^{\text{th}}$ feature and response variable respectively. Let $\ell$ and $R$ be the loss function and regularizer, $\boldsymbol{\beta}$ denote the unknown weights, and $\lambda$ be a regularization parameter. Finding the optimal choice of $\lambda$ is a challenging problem in high-dimensional regimes where both $n$ and $p$ are large. We propose two frameworks to obtain a computationally efficient approximation ALO of the leave-one-out cross validation (LOOCV) risk for nonsmooth losses and regularizers. Our two frameworks are based on the primal and dual formulations of (1). We prove the equivalence of the two approaches under smoothness conditions. This equivalence enables us to justify the accuracy of both methods under such conditions. We use our approaches to obtain a risk estimate for several standard problems, including generalized LASSO, nuclear norm regularization, and support vector machines. We empirically demonstrate the effectiveness of our results for non-differentiable cases.
Predicting Infant Motor Development Status using Day Long Movement Data from Wearable Sensors
Goodfellow, David, Zhi, Ruoyu, Funke, Rebecca, Pulido, Jose Carlos, Mataric, Maja, Smith, Beth A.
Infants with a variety of complications at or before birth are classified as being at risk for developmental delays (AR). As they grow older, they are followed by healthcare providers in an effort to discern whether they are on a typical or impaired developmental trajectory. Often, it is difficult to make an accurate determination early in infancy as infants with typical development (TD) display high variability in their developmental trajectories both in content and timing. Studies have shown that spontaneous movements have the potential to differentiate typical and atypical trajectories early in life using sensors and kinematic analysis systems. In this study, machine learning classification algorithms are used to take inertial movement from wearable sensors placed on an infant for a day and predict if the infant is AR or TD, thus further establishing the connection between early spontaneous movement and developmental trajectory.
Certifying Global Optimality of Graph Cuts via Semidefinite Relaxation: A Performance Guarantee for Spectral Clustering
Ling, Shuyang, Strohmer, Thomas
Spectral clustering has become one of the most widely used clustering techniques when the structure of the individual clusters is non-convex or highly anisotropic. Yet, despite its immense popularity, there exists fairly little theory about performance guarantees for spectral clustering. This issue is partly due to the fact that spectral clustering typically involves two steps which complicated its theoretical analysis: first, the eigenvectors of the associated graph Laplacian are used to embed the dataset, and second, k-means clustering algorithm is applied to the embedded dataset to get the labels. This paper is devoted to the theoretical foundations of spectral clustering and graph cuts. We consider a convex relaxation of graph cuts, namely ratio cuts and normalized cuts, that makes the usual two-step approach of spectral clustering obsolete and at the same time gives rise to a rigorous theoretical analysis of graph cuts and spectral clustering. We derive deterministic bounds for successful spectral clustering via a spectral proximity condition that naturally depends on the algebraic connectivity of each cluster and the inter-cluster connectivity. Moreover, we demonstrate by means of some popular examples that our bounds can achieve near-optimality. Our findings are also fundamental for the theoretical understanding of kernel k-means. Numerical simulations confirm and complement our analysis.
Mirror descent in saddle-point problems: Going the extra (gradient) mile
Mertikopoulos, Panayotis, Zenati, Houssam, Lecouat, Bruno, Foo, Chuan-Sheng, Chandrasekhar, Vijay, Piliouras, Georgios
Owing to their connection with generative adversarial networks (GANs), saddle-point problems have recently attracted considerable interest in machine learning and beyond. By necessity, most theoretical guarantees revolve around convex-concave problems; however, making theoretical inroads towards efficient GAN training crucially depends on moving beyond this classic framework. To make piecemeal progress along these lines, we analyze the widely used mirror descent (MD) method in a class of non-monotone problems - called coherent - whose solutions coincide with those of a naturally associated variational inequality. Our first result is that, under strict coherence (a condition satisfied by all strictly convex-concave problems), MD methods converge globally; however, they may fail to converge even in simple, bilinear models. To mitigate this deficiency, we add on an "extra-gradient" step which we show stabilizes MD methods by looking ahead and using a "future gradient". These theoretical results are subsequently validated by numerical experiments in GANs.
Gradient Hyperalignment for multi-subject fMRI data alignment
Xu, Tonglin, Yousefnezhad, Muhammad, Zhang, Daoqiang
Multi-subject fMRI data analysis is an interesting and challenging problem in human brain decoding studies. The inherent anatomical and functional variability across subjects make it necessary to do both anatomical and functional alignment before classification analysis. Besides, when it comes to big data, time complexity becomes a problem that cannot be ignored. This paper proposes Gradient Hyperalignment (Gradient-HA) as a gradient-based functional alignment method that is suitable for multi-subject fMRI datasets with large amounts of samples and voxels. The advantage of Gradient-HA is that it can solve independence and high dimension problems by using Independent Component Analysis (ICA) and Stochastic Gradient Ascent (SGA). Validation using multi-classification tasks on big data demonstrates that Gradient-HA method has less time complexity and better or comparable performance compared with other state-of-the-art functional alignment methods.
A Supervised Geometry-Aware Mapping Approach for Classification of Hyperspectral Images
Mohanty, Ramanarayan, Happy, S L, Routray, Aurobinda
The multi-path scattering of light within a pixel [1], bidirectional reflectance distribution [2], and the heterogeneity of sub-pixel constituents [3] are the major concerns in the hyperspectral (HS) data classification. These nonlinearity properties naturally place the HS data on a non-euclidean space. Handling these high dimensional redundant data in a non-euclidean space is one of the major bottlenecks in HS data analysis. Typically, HS classification consists of dimensionality reduction (DR) and subsequent classification operation. The popular DR methods such as principal component analysis (PCA) [4] and linear discriminant analysis (LDA) [5] are linear and operate on Euclidean structures. These linear DR methods skip the curved nonlinear structures of the HS data. On the other hand, manifold learning helps in recovering compact, meaningful low dimensional structures from those complex high dimensional data from a non-euclidean space. The manifold learning methods consider the real world high dimensional data to be generated with a few degrees of freedom [6]. This leads to the projection of the data into lower dimensional space while preserving their underlying geometrical structure [7].
Temporal graph-based clustering for historical record linkage
Nanayakkara, Charini, Christen, Peter, Ranbaduge, Thilina
Research in the social sciences is increasingly based on large and complex data collections, where individual data sets from different domains are linked and integrated to allow advanced analytics. A popular type of data used in such a context are historical censuses, as well as birth, death, and marriage certificates. Individually, such data sets however limit the types of studies that can be conducted. Specifically, it is impossible to track individuals, families, or households over time. Once such data sets are linked and family trees spanning several decades are available it is possible to, for example, investigate how education, health, mobility, employment, and social status influence each other and the lives of people over two or even more generations. A major challenge is however the accurate linkage of historical data sets which is due to data quality and commonly also the lack of ground truth data being available. Unsupervised techniques need to be employed, which can be based on similarity graphs generated by comparing individual records. In this paper we present initial results from clustering birth records from Scotland where we aim to identify all births of the same mother and group siblings into clusters. We extend an existing clustering technique for record linkage by incorporating temporal constraints that must hold between births by the same mother, and propose a novel greedy temporal clustering technique. Experimental results show improvements over non-temporary approaches, however further work is needed to obtain links of high quality.
3D Steerable CNNs: Learning Rotationally Equivariant Features in Volumetric Data
Weiler, Maurice, Geiger, Mario, Welling, Max, Boomsma, Wouter, Cohen, Taco
We present a convolutional network that is equivariant to rigid body motions. The model uses scalar-, vector-, and tensor fields over 3D Euclidean space to represent data, and equivariant convolutions to map between such representations. These SE(3)-equivariant convolutions utilize kernels which are parameterized as a linear combination of a complete steerable kernel basis, which is derived in this paper. We prove that equivariant convolutions are the most general equivariant linear maps between fields over R^3. Our experimental results confirm the effectiveness of 3D Steerable CNNs for the problem of amino acid propensity prediction and protein structure classification, both of which have inherent SE(3) symmetry.