Performance Analysis
Bootstrap Bias Corrected Cross Validation applied to Super Learning
Mnich, Krzysztof, Golińska, Agnieszka Kitlas, Polewko-Klim, Aneta, Rudnicki, Witold R.
Super learner algorithm can be applied to combine results of multiple base learners to improve quality of predictions. The default method for verification of super learner results is by nested cross validation. It has been proposed by Tsamardinos et al., that nested cross validation can be replaced by resampling for tuning hyper-parameters of the learning algorithms. We apply this idea to verification of super learner and compare with other verification methods, including nested cross validation. Tests were performed on artificial data sets of diverse size and on seven real, biomedical data sets. The resampling method, called Bootstrap Bias Correction, proved to be a reasonably precise and very cost-efficient alternative for nested cross validation.
Detecting COVID-19 in X-ray images with Keras, TensorFlow, and Deep Learning - PyImageSearch
In this tutorial, you will learn how to automatically detect COVID-19 in a hand-created X-ray image dataset using Keras, TensorFlow, and Deep Learning. Like most people in the world right now, I'm genuinely concerned about COVID-19. I find myself constantly analyzing my personal health and wondering if/when I will contract it. At first, I didn't think much of it -- I have pollen allergies and due to the warm weather on the eastern coast of the United States, spring has come early this year. My allergies were likely just acting up. But my symptoms didn't improve throughout the day. I'm actually sitting here, writing the this tutorial, with a thermometer in my mouth; and glancing down I see that it reads 99.4 Fahrenheit. My body runs a bit cooler than most, typically in the 97.4 F range.
An Automatic Attribute Based Access Control Policy Extraction from Access Logs
Karimi, Leila, Aldairi, Maryam, Joshi, James, Abdelhakim, Mai
With the rapid advances in computing and information technologies, traditional access control models have become inadequate in terms of capturing fine-grained, and expressive security requirements of newly emerging applications. An attribute-based access control (ABAC) model provides a more flexible approach for addressing the authorization needs of complex and dynamic systems. While organizations are interested in employing newer authorization models, migrating to such models pose as a significant challenge. Many large-scale businesses need to grant authorization to their user populations that are potentially distributed across disparate and heterogeneous computing environments. Each of these computing environments may have its own access control model. The manual development of a single policy framework for an entire organization is tedious, costly, and error-prone. In this paper, we present a methodology for automatically learning ABAC policy rules from access logs of a system to simplify the policy development process. The proposed approach employs an unsupervised learning-based algorithm for detecting patterns in access logs and extracting ABAC authorization rules from these patterns. In addition, we present two policy improvement algorithms, including rule pruning and policy refinement algorithms to generate a higher quality mined policy. Finally, we implement a prototype of the proposed approach to demonstrate its feasibility.
CARPAL: Confidence-Aware Intent Recognition for Parallel Autonomy
Huang, Xin, McGill, Stephen G., DeCastro, Jonathan A., Williams, Brian C., Fletcher, Luke, Leonard, John J., Rosman, Guy
Predicting the behavior of road agents is a difficult and crucial task for both advanced driver assistance and autonomous driving systems. Traditional confidence measures for this important task often ignore the way predicted trajectories affect downstream decisions and their utilities. In this paper we devise a novel neural network regressor to estimate the utility distribution given the predictions. Based on reasonable assumptions on the utility function, we establish a decision criterion that takes into account the role of prediction in decision making. We train our real-time regressor along with a human driver intent predictor and use it in shared autonomy scenarios where decisions depend on the prediction confidence. We test our system on a realistic urban driving dataset, present the advantage of the resulting system in terms of recall and fall-out performance compared to baseline methods, and demonstrate its effectiveness in intervention and warning use cases.
Adversarial Transferability in Wearable Sensor Systems
Sah, Ramesh Kumar, Ghasemzadeh, Hassan
Machine learning has increasingly become the most used approach for inference and decision making in wearable sensor systems. However, recent studies have found that machine learning systems are easily fooled by the addition of adversarial perturbation to their inputs. What is more interesting is that the adversarial examples generated for one machine learning system can also degrade the performance of another. This property of adversarial examples is called transferability. In this work, we take the first strides in studying adversarial transferability in wearable sensor systems, from the following perspectives: 1) Transferability between machine learning models, 2) Transferability across subjects, 3) Transferability across sensor locations, and 4) Transferability across datasets. With Human Activity Recognition (HAR) as an example sensor system, we found strong untargeted transferability in all cases of transferability. Specifically, gradient-based attacks were able to achieve higher misclassification rates compared to non-gradient attacks. The misclassification rate of untargeted adversarial examples ranged from 20% to 98%. For targeted transferability between machine learning models, the success rate of adversarial examples was 100% for iterative attack methods. However, the success rate for other types of targeted transferability ranged from 20% to 0%. Our findings strongly suggest that adversarial transferability has serious consequences not only in sensor systems but also across the broad spectrum of ubiquitous computing.
ParKCa: Causal Inference with Partially Known Causes
Causal Inference methods based on observational data are an alternative for applications where collecting the counterfactual data or realizing a more standard experiment is not possible. In this work, our goal is to combine several observational causal inference methods to learn new causes in applications where some causes are well known. We validate the proposed method on The Cancer Genome Atlas (TCGA) dataset to identify genes that potentially cause metastasis.
Key Phrase Classification in Complex Assignments
Complex assignments typically consist of open-ended questions with large and diverse content in the context of both classroom and online graduate programs. With the sheer scale of these programs comes a variety of problems in peer and expert feedback, including rogue reviews. As such with the hope of identifying important contents needed for the review, in this work we present a very first work on key phrase classification with a detailed empirical study on traditional and most recent language modeling approaches. From this study, we find that the task of classification of key phrases is ambiguous at a human level producing Cohen's kappa of 0.77 on a new data set. Both pretrained language models and simple TFIDF SVM classifiers produce similar results with a former producing average of 0.6 F1 higher than the latter. We finally derive practical advice from our extensive empirical and model interpretability results for those interested in key phrase classification from educational reports in the future.
AutoCogniSys: IoT Assisted Context-Aware Automatic Cognitive Health Assessment
Alam, Mohammad Arif Ul, Roy, Nirmalya, Holmes, Sarah, Gangopadhyay, Aryya, Galik, Elizabeth
Cognitive impairment has become epidemic in older adult population. The recent advent of tiny wearable and ambient devices, a.k.a Internet of Things (IoT) provides ample platforms for continuous functional and cognitive health assessment of older adults. In this paper, we design, implement and evaluate AutoCogniSys, a context-aware automated cognitive health assessment system, combining the sensing powers of wearable physiological (Electrodermal Activity, Photoplethysmography) and physical (Accelerometer, Object) sensors in conjunction with ambient sensors. We design appropriate signal processing and machine learning techniques, and develop an automatic cognitive health assessment system in a natural older adults living environment. We validate our approaches using two datasets: (i) a naturalistic sensor data streams related to Activities of Daily Living and mental arousal of 22 older adults recruited in a retirement community center, individually living in their own apartments using a customized inexpensive IoT system (IRB #HP-00064387) and (ii) a publicly available dataset for emotion detection. The performance of AutoCogniSys attests max. 93\% of accuracy in assessing cognitive health of older adults.
Why understanding your fraud false-positive rate is key to growing your business
'Ecommerce businesses have a problem - one that causes lost customer revenue, yet has been historically nearly impossible to solve' Geoff Huang, VP of Product at Sift The problem stems from the inability to know their false-positive rate, which is the percentage of orders from legitimate customers that are mistakenly blocked as fraud. According to a survey conducted by CNP, 42% of ecommerce merchants don't know their false-positive rate (also known as customer insult rate). That is a startling statistic--nearly half of online sellers have no visibility into the number of good orders they inadvertently block or the subsequent revenue lost from those orders. And the news, unfortunately, doesn't get much better. Sift polled 1,000 adult consumers and found roughly 25% of insulted online shoppers--those who were falsely declined--will take their business to a competitor.
DriftSurf: A Risk-competitive Learning Algorithm under Concept Drift
Tahmasbi, Ashraf, Jothimurugesan, Ellango, Tirthapura, Srikanta, Gibbons, Phillip B.
When learning from streaming data, a change in the data distribution, also known as concept drift, can render a previously-learned model inaccurate and require training a new model. We present an adaptive learning algorithm that extends previous drift-detection-based methods by incorporating drift detection into a broader stable-state/reactive-state process. The advantage of our approach is that we can use aggressive drift detection in the stable state to achieve a high detection rate, but mitigate the false positive rate of standalone drift detection via a reactive state that reacts quickly to true drifts while eliminating most false positives. The algorithm is generic in its base learner and can be applied across a variety of supervised learning problems. Our theoretical analysis shows that the risk of the algorithm is competitive to an algorithm with oracle knowledge of when (abrupt) drifts occur. Experiments on synthetic and real datasets with concept drifts confirm our theoretical analysis.