Deep Learning
Human Activity Analysis and Recognition from Smartphones using Machine Learning Techniques
Rabbi, Jakaria, Fuad, Md. Tahmid Hasan, Awal, Md. Abdul
Human Activity Recognition (HAR) is considered a valuable research topic in the last few decades. Different types of machine learning models are used for this purpose, and this is a part of analyzing human behavior through machines. It is not a trivial task to analyze the data from wearable sensors for complex and high dimensions. Nowadays, researchers mostly use smartphones or smart home sensors to capture these data. In our paper, we analyze these data using machine learning models to recognize human activities, which are now widely used for many purposes such as physical and mental health monitoring. We apply different machine learning models and compare performances. We use Logistic Regression (LR) as the benchmark model for its simplicity and excellent performance on a dataset, and to compare, we take Decision Tree (DT), Support Vector Machine (SVM), Random Forest (RF), and Artificial Neural Network (ANN). Additionally, we select the best set of parameters for each model by grid search. We use the HAR dataset from the UCI Machine Learning Repository as a standard dataset to train and test the models. Throughout the analysis, we can see that the Support Vector Machine performed (average accuracy 96.33%) far better than the other methods. We also prove that the results are statistically significant by employing statistical significance test methods.
EnergyVis: Interactively Tracking and Exploring Energy Consumption for ML Models
Shaikh, Omar, Saad-Falcon, Jon, Wright, Austin P, Das, Nilaksh, Freitas, Scott, Asensio, Omar Isaac, Chau, Duen Horng
The advent of larger machine learning (ML) models have improved state-of-the-art (SOTA) performance in various modeling tasks, ranging from computer vision to natural language. As ML models continue increasing in size, so does their respective energy consumption and computational requirements. However, the methods for tracking, reporting, and comparing energy consumption remain limited. We presentEnergyVis, an interactive energy consumption tracker for ML models. Consisting of multiple coordinated views, EnergyVis enables researchers to interactively track, visualize and compare model energy consumption across key energy consumption and carbon footprint metrics (kWh and CO2), helping users explore alternative deployment locations and hardware that may reduce carbon footprints. EnergyVis aims to raise awareness concerning computational sustainability by interactively highlighting excessive energy usage during model training; and by providing alternative training options to reduce energy usage.
User profile-driven large-scale multi-agent learning from demonstration in federated human-robot collaborative environments
Papadopoulos, Georgios Th., Leonidis, Asterios, Antona, Margherita, Stephanidis, Constantine
Learning from Demonstration (LfD) has been established as the dominant paradigm for efficiently transferring skills from human teachers to robots. In this context, the Federated Learning (FL) conceptualization has very recently been introduced for developing large-scale human-robot collaborative environments, targeting to robustly address, among others, the critical challenges of multi-agent learning and long-term autonomy. In the current work, the latter scheme is further extended and enhanced, by designing and integrating a novel user profile formulation for providing a fine-grained representation of the exhibited human behavior, adopting a Deep Learning (DL)-based formalism. In particular, a hierarchically organized set of key information sources is considered, including: a) User attributes (e.g. demographic, anthropomorphic, educational, etc.), b) User state (e.g. fatigue detection, stress detection, emotion recognition, etc.) and c) Psychophysiological measurements (e.g. gaze, electrodermal activity, heart rate, etc.) related data. Then, a combination of Long Short-Term Memory (LSTM) and stacked autoencoders, with appropriately defined neural network architectures, is employed for the modelling step. The overall designed scheme enables both short- and long-term analysis/interpretation of the human behavior (as observed during the feedback capturing sessions), so as to adaptively adjust the importance of the collected feedback samples when aggregating information originating from the same and different human teachers, respectively.
Temporal Memory Relation Network for Workflow Recognition from Surgical Video
Jin, Yueming, Long, Yonghao, Chen, Cheng, Zhao, Zixu, Dou, Qi, Heng, Pheng-Ann
Automatic surgical workflow recognition is a key component for developing context-aware computer-assisted systems in the operating theatre. Previous works either jointly modeled the spatial features with short fixed-range temporal information, or separately learned visual and long temporal cues. In this paper, we propose a novel end-to-end temporal memory relation network (TMRNet) for relating long-range and multi-scale temporal patterns to augment the present features. We establish a long-range memory bank to serve as a memory cell storing the rich supportive information. Through our designed temporal variation layer, the supportive cues are further enhanced by multi-scale temporal-only convolutions. To effectively incorporate the two types of cues without disturbing the joint learning of spatio-temporal features, we introduce a non-local bank operator to attentively relate the past to the present. In this regard, our TMRNet enables the current feature to view the long-range temporal dependency, as well as tolerate complex temporal extents. We have extensively validated our approach on two benchmark surgical video datasets, M2CAI challenge dataset and Cholec80 dataset. Experimental results demonstrate the outstanding performance of our method, consistently exceeding the state-of-the-art methods by a large margin (e.g., 67.0% v.s. 78.9% Jaccard on Cholec80 dataset).
Exploring Edge TPU for Network Intrusion Detection in IoT
Hosseininoorbin, Seyedehfaezeh, Layeghy, Siamak, Sarhan, Mohanad, Jurdak, Raja, Portmann, Marius
This paper explores Google's Edge TPU for implementing a practical network intrusion detection system (NIDS) at the edge of IoT, based on a deep learning approach. While there are a significant number of related works that explore machine learning based NIDS for the IoT edge, they generally do not consider the issue of the required computational and energy resources. The focus of this paper is the exploration of deep learning-based NIDS at the edge of IoT, and in particular the computational and energy efficiency. In particular, the paper studies Google's Edge TPU as a hardware platform, and considers the following three key metrics: computation (inference) time, energy efficiency and the traffic classification performance. Various scaled model sizes of two major deep neural network architectures are used to investigate these three metrics. The performance of the Edge TPU-based implementation is compared with that of an energy efficient embedded CPU (ARM Cortex A53). Our experimental evaluation shows some unexpected results, such as the fact that the CPU significantly outperforms the Edge TPU for small model sizes.
Model-Contrastive Federated Learning
Li, Qinbin, He, Bingsheng, Song, Dawn
Federated learning enables multiple parties to collaboratively train a machine learning model without communicating their local data. A key challenge in federated learning is to handle the heterogeneity of local data distribution across parties. Although many studies have been proposed to address this challenge, we find that they fail to achieve high performance in image datasets with deep learning models. In this paper, we propose MOON: model-contrastive federated learning. MOON is a simple and effective federated learning framework. The key idea of MOON is to utilize the similarity between model representations to correct the local training of individual parties, i.e., conducting contrastive learning in model-level. Our extensive experiments show that MOON significantly outperforms the other state-of-the-art federated learning algorithms on various image classification tasks.
Convolutional Neural Networks for Sleep Stage Scoring on a Two-Channel EEG Signal
Fernandez-Blanco, Enrique, Rivero, Daniel, Pazos, Alejandro
Among the essential body functions like breathing, eating or drinking, sleeping is probably the most problematic one nowadays. According to the US government through its Centers for Control of Disease and Prevention (CDC), about 9 million citizens have frequent problems to develop good quality sleep and end up resorting to sleeping pills (Ford et al. 2014). In parallel, recent studies (Stranges et al. 2012; Chong et al. 2013) have estimated that at least 15% of adult population might have some kind of sleeping problem or poor-quality sleep as a result of a number of issues. Moreover, the World Health Organization (WHO) (2015) claimed that a good quality sleep was one of the most important factors for good health while sleeping problems were directly related to other diseases, including depression, stress or early cardiac diseases. As a consequence, new units focused on the study and treatment of sleeping problems have been created in hospitals all over the world. The physicians in these units have as their main tool for their work the records obtained during their patients' sleep. These records, called polysomnography (PSG), may include a great variety of signals such as Electrocardiograms, Electroencephalograms, respiratory signals or movement records. Among these signals, the most important one is the Electroencephalogram (EEG) because it is the most reliable to determine the sleep stage a patient is in. The interpretation of an EEG is a highly time-consuming activity (Akben and Alkan 2016), which usually requires a specialist and it is deeply dependent on the expert's expertise.
MT3: Meta Test-Time Training for Self-Supervised Test-Time Adaption
Bartler, Alexander, Bรผhler, Andre, Wiewel, Felix, Dรถbler, Mario, Yang, Bin
An unresolved problem in Deep Learning is the ability of neural networks to cope with domain shifts during test-time, imposed by commonly fixing network parameters after training. Our proposed method Meta Test-Time Training (MT3), however, breaks this paradigm and enables adaption at test-time. We combine meta-learning, self-supervision and test-time training to learn to adapt to unseen test distributions. By minimizing the self-supervised loss, we learn task-specific model parameters for different tasks. A meta-model is optimized such that its adaption to the different task-specific models leads to higher performance on those tasks. During test-time a single unlabeled image is sufficient to adapt the meta-model parameters. This is achieved by minimizing only the self-supervised loss component resulting in a better prediction for that image. Our approach significantly improves the state-of-the-art results on the CIFAR-10-Corrupted image classification benchmark. Our implementation is available on GitHub.
Prediction of Landfall Intensity, Location, and Time of a Tropical Cyclone
Kumar, Sandeep, Biswas, Koushik, Pandey, Ashish Kumar
TC is characterised by warm core, and a low and availability of huge data, new models using Artificial pressure system with a large vortex in the atmosphere. TC Neural Networks (ANNs) have been increasingly used to brings strong winds, heavy precipitation and high tides in forecast track and intensity of cyclones (Leroux et al. 2018; coastal areas and resulted in huge economic and human loss. Alemany et al. 2018; Giffard-Roisin et al. 2020; Moradi Kordmahalleh, Over the years, many destructive TCs have originated in the Gorji Sefidmazgi, and Homaifar 2016). North Indian Ocean (NIO), consisting of the Bay of Bengal The most important prediction about a TC is its arrival at and the Arabian Sea. In 2008, Nargis, one of the disastrous land, known as landfall of a cyclone. The accurate prediction TC in recent times, originated in the Bay of Bengal and resulted about the location and time of the landfall, and intensity of in 13,800 casualties alone in Myanmar and caused the cyclone at the landfall will hugely help authorities to take US$15.4 billion economic loss (Fritz et al. 2009). In 2018, preventive measures and reduce material and human loss. In Fani cyclone caused 89 causalities in India and Bangladesh, this work, we attempt to predict intensity, location, and time and US$9.1 billion economic loss (Kumar, Lal, and Kumar of the landfall of a TC at any instance of time during the 2020).
Predicting Landfall's Location and Time of a Tropical Cyclone Using Reanalysis Data
Kumar, Sandeep, Biswas, Koushik, Pandey, Ashish Kumar
Landfall of a tropical cyclone is the event when it moves over the land after crossing the coast of the ocean. It is important to know the characteristics of the landfall in terms of location and time, well advance in time to take preventive measures timely. In this article, we develop a deep learning model based on the combination of a Convolutional Neural network and a Long Short-Term memory network to predict the landfall's location and time of a tropical cyclone in six ocean basins of the world with high accuracy. We have used high-resolution spacial reanalysis data, ERA5, maintained by European Center for Medium-Range Weather Forecasting (ECMWF). The model takes any 9 hours, 15 hours, or 21 hours of data, during the progress of a tropical cyclone and predicts its landfall's location in terms of latitude and longitude and time in hours. For 21 hours of data, we achieve mean absolute error for landfall's location prediction in the range of 66.18 - 158.92 kilometers and for landfall's time prediction in the range of 4.71 - 8.20 hours across all six ocean basins. The model can be trained in just 30 to 45 minutes (based on ocean basin) and can predict the landfall's location and time in a few seconds, which makes it suitable for real time prediction.