Support Vector Machines
Never a Dull Moment: Distributional Properties as a Baseline for Time-Series Classification
Henderson, Trent, Bryant, Annie G., Fulcher, Ben D.
The variety of complex algorithmic approaches for tackling time-series classification problems has grown considerably over the past decades, including the development of sophisticated but challenging-to-interpret deep-learning-based methods. But without comparison to simpler methods it can be difficult to determine when such complexity is required to obtain strong performance on a given problem. Here we evaluate the performance of an extremely simple classification approach -- a linear classifier in the space of two simple features that ignore the sequential ordering of the data: the mean and standard deviation of time-series values. Across a large repository of 128 univariate time-series classification problems, this simple distributional moment-based approach outperformed chance on 69 problems, and reached 100% accuracy on two problems. With a neuroimaging time-series case study, we find that a simple linear model based on the mean and standard deviation performs better at classifying individuals with schizophrenia than a model that additionally includes features of the time-series dynamics. Comparing the performance of simple distributional features of a time series provides important context for interpreting the performance of complex time-series classification models, which may not always be required to obtain high accuracy.
An ADMM Solver for the MKL-$L_{0/1}$-SVM
We formulate the Multiple Kernel Learning (abbreviated as MKL) problem for the support vector machine with the infamous $(0,1)$-loss function. Some first-order optimality conditions are given and then exploited to develop a fast ADMM solver for the nonconvex and nonsmooth optimization problem. A simple numerical experiment on synthetic planar data shows that our MKL-$L_{0/1}$-SVM framework could be promising.
Domain Knowledge integrated for Blast Furnace Classifier Design
Chen, Shaohan, Fan, Di, Gao, Chuanhou
Blast furnace modeling and control is one of the important problems in the industrial field, and the black-box model is an effective mean to describe the complex blast furnace system. In practice, there are often different learning targets, such as safety and energy saving in industrial applications, depending on the application. For this reason, this paper proposes a framework to design a domain knowledge integrated classification model that yields a classifier for industrial application. Our knowledge incorporated learning scheme allows the users to create a classifier that identifies "important samples" (whose misclassifications can lead to severe consequences) more correctly, while keeping the proper precision of classifying the remaining samples. The effectiveness of the proposed method has been verified by two real blast furnace datasets, which guides the operators to utilize their prior experience for controlling the blast furnace systems better.
Attention is Not Always What You Need: Towards Efficient Classification of Domain-Specific Text
Wahba, Yasmen, Madhavji, Nazim, Steinbacher, John
For large-scale IT corpora with hundreds of classes organized in a hierarchy, the task of accurate classification of classes at the higher level in the hierarchies is crucial to avoid errors propagating to the lower levels. In the business world, an efficient and explainable ML model is preferred over an expensive black-box model, especially if the performance increase is marginal. A current trend in the Natural Language Processing (NLP) community is towards employing huge pre-trained language models (PLMs) or what is known as self-attention models (e.g., BERT) for almost any kind of NLP task (e.g., question-answering, sentiment analysis, text classification). Despite the widespread use of PLMs and the impressive performance in a broad range of NLP tasks, there is a lack of a clear and well-justified need to as why these models are being employed for domain-specific text classification (TC) tasks, given the monosemic nature of specialized words (i.e., jargon) found in domain-specific text which renders the purpose of contextualized embeddings (e.g., PLMs) futile. In this paper, we compare the accuracies of some state-of-the-art (SOTA) models reported in the literature against a Linear SVM classifier and TFIDF vectorization model on three TC datasets. Results show a comparable performance for the LinearSVM. The findings of this study show that for domain-specific TC tasks, a linear model can provide a comparable, cheap, reproducible, and interpretable alternative to attention-based models.
Using Connected Vehicle Trajectory Data to Evaluate the Effects of Speeding
Ugan, Jorge, Abdel-Aty, Mohamed, Islam, Zubayer
Speeding has been and continues to be a major contributing factor to traffic fatalities. Various transportation agencies have proposed speed management strategies to reduce the amount of speeding on arterials. While there have been various studies done on the analysis of speeding proportions above the speed limit, few studies have considered the effect on the individual's journey. Many studies utilized speed data from detectors, which is limited in that there is no information of the route that the driver took. This study aims to explore the effects of various roadway features an individual experiences for a given journey on speeding proportions. Connected vehicle trajectory data was utilized to identify the path that a driver took, along with the vehicle related variables. The level of speeding proportion is predicted using multiple learning models. The model with the best performance, Extreme Gradient Boosting, achieved an accuracy of 0.756. The proposed model can be used to understand how the environment and vehicle's path effects the drivers' speeding behavior, as well as predict the areas with high levels of speeding proportions. The results suggested that features related to an individual driver's trip, i.e., total travel time, has a significant contribution towards speeding. Features that are related to the environment of the individual driver's trip, i.e., proportion of residential area, also had a significant effect on reducing speeding proportions. It is expected that the findings could help inform transportation agencies more on the factors related to speeding for an individual driver's trip.
On the Connection between $L_p$ and Risk Consistency and its Implications on Regularized Kernel Methods
As a predictor's quality is often assessed by means of its risk, it is natural to regard risk consistency as a desirable property of learning methods, and many such methods have indeed been shown to be risk consistent. The first aim of this paper is to establish the close connection between risk consistency and $L_p$-consistency for a considerably wider class of loss functions than has been done before. The attempt to transfer this connection to shifted loss functions surprisingly reveals that this shift does not reduce the assumptions needed on the underlying probability measure to the same extent as it does for many other results. The results are applied to regularized kernel methods such as support vector machines.
Attention-based Learning for Sleep Apnea and Limb Movement Detection using Wi-Fi CSI Signals
Chang, Chi-Che, Hsiao, An-Hung, Shen, Li-Hsiang, Feng, Kai-Ten, Chen, Chia-Yu
Wi-Fi channel state information (CSI) has become a promising solution for non-invasive breathing and body motion monitoring during sleep. Sleep disorders of apnea and periodic limb movement disorder (PLMD) are often unconscious and fatal. The existing researches detect abnormal sleep disorders in impractically controlled environments. Moreover, it leads to compelling challenges to classify complex macro- and micro-scales of sleep movements as well as entangled similar waveforms of cases of apnea and PLMD. In this paper, we propose the attention-based learning for sleep apnea and limb movement detection (ALESAL) system that can jointly detect sleep apnea and PLMD under different sleep postures across a variety of patients. ALESAL contains antenna-pair and time attention mechanisms for mitigating the impact of modest antenna pairs and emphasizing the duration of interest, respectively. Performance results show that our proposed ALESAL system can achieve a weighted F1-score of 84.33, outperforming the other existing non-attention based methods of support vector machine and deep multilayer perceptron.
Utilizing Network Properties to Detect Erroneous Inputs
Gorbett, Matt, Blanchard, Nathaniel
Neural networks are vulnerable to a wide range of erroneous inputs such as adversarial, corrupted, out-of-distribution, and misclassified examples. In this work, we train a linear SVM classifier to detect these four types of erroneous data using hidden and softmax feature vectors of pre-trained neural networks. Our results indicate that these faulty data types generally exhibit linearly separable activation properties from correct examples, giving us the ability to reject bad inputs with no extra training or overhead. We experimentally validate our findings across a diverse range of datasets, domains, pre-trained models, and adversarial attacks.
Machine Learning Detects Multiplicity of the First Stars in Stellar Archaeology Data - IOPscience
In unveiling the nature of the first stars, the main astronomical clue is the elemental compositions of the second generation of stars, observed as extremely metal-poor (EMP) stars, in the Milky Way. However, no observational constraint was available on their multiplicity, which is crucial for understanding early phases of galaxy formation. We develop a new data-driven method to classify observed EMP stars into mono- or multi-enriched stars with support vector machines. We also use our own nucleosynthesis yields of core-collapse supernovae with mixing fallback that can explain many of the observed EMP stars. Our method predicts, for the first time, that 31.8%
Extracting Physical Rehabilitation Exercise Information from Clinical Notes: a Comparison of Rule-Based and Machine Learning Natural Language Processing Techniques
Shaffran, Stephen W., Gao, Fengyi, Denny, Parker E., Aldhahwani, Bayan M., Bove, Allyn, Visweswaran, Shyam, Wang, Yanshan
However, physical therapy procedures are typically described in unstructured clinical notes, meaning that simple data extraction methods such as database queries cannot be applied to obtain sufficient information. Additionally, the language used to describe these procedures can differ between clinicians, cites, and times. A more advanced natural language processing (NLP) algorithm is required to extract this information from clinical notes, but such a method has not yet been developed for this application. In this paper we devise and compare several approaches to extracting information about therapeutic procedures for physical rehabilitation, both for the purpose of emulating a manual annotation process using named entity recognition (NER) and categorizing descriptions of therapeutic procedures using multi label sequence classification. Using a set of manually annotated notes as a gold standard reference, we evaluated the performance of a rule-based algorithm using the MedTagger software, and several machine learning approaches such as logistic regression (LR) and support vector machines (SVM). Methods Data Collection We identified a cohort of patients diagnosed with stroke between January 1st, 2016 and December 31st, 2016 at UPMC. For these patients, we extracted clinical encounter notes created between January 1st, 2016 and December 31st, 2018 from the institutional data warehouse. The study was approved by the University of Pittsburgh's Institutional Review Board (IRB #21040204).