Deep Learning
Probabilistic multivariate electricity price forecasting using implicit generative ensemble post-processing
The reliable estimation of forecast uncertainties is crucial for risk-sensitive optimal decision making. In this paper, we propose implicit generative ensemble post-processing, a novel framework for multivariate probabilistic electricity price forecasting. We use a likelihood-free implicit generative model based on an ensemble of point forecasting models to generate multivariate electricity price scenarios with a coherent dependency structure as a representation of the joint predictive distribution. Our ensemble post-processing method outperforms well-established model combination benchmarks. This is demonstrated on a data set from the German day-ahead market. As our method works on top of an ensemble of domain-specific expert models, it can readily be deployed to other forecasting tasks.
General-Purpose User Embeddings based on Mobile App Usage
Zhang, Junqi, Bai, Bing, Lin, Ye, Liang, Jian, Bai, Kun, Wang, Fei
In this paper, we report our recent practice at Tencent for user modeling based on mobile app usage. User behaviors on mobile app usage, including retention, installation, and uninstallation, can be a good indicator for both long-term and short-term interests of users. For example, if a user installs Snapseed recently, she might have a growing interest in photographing. Such information is valuable for numerous downstream applications, including advertising, recommendations, etc. Traditionally, user modeling from mobile app usage heavily relies on handcrafted feature engineering, which requires onerous human work for different downstream applications, and could be sub-optimal without domain experts. However, automatic user modeling based on mobile app usage faces unique challenges, including (1) retention, installation, and uninstallation are heterogeneous but need to be modeled collectively, (2) user behaviors are distributed unevenly over time, and (3) many long-tailed apps suffer from serious sparsity. In this paper, we present a tailored AutoEncoder-coupled Transformer Network (AETN), by which we overcome these challenges and achieve the goals of reducing manual efforts and boosting performance. We have deployed the model at Tencent, and both online/offline experiments from multiple domains of downstream applications have demonstrated the effectiveness of the output user embeddings.
Fast and Effective Robustness Certification for Recurrent Neural Networks
Ryou, Wonryong, Chen, Jiayu, Balunovic, Mislav, Singh, Gagandeep, Dan, Andrei, Vechev, Martin
We present a precise and scalable verifier for recurrent neural networks, called R2. The verifier is based on two key ideas: (i) a method to compute tight linear convex relaxations of a recurrent update function via sampling and optimization, and (ii) a technique to optimize convex combinations of multiple bounds for each neuron instead of a single bound as previously done. Using R2, we present the first study of certifying a non-trivial use case of recurrent neural networks, namely speech classification. This required us to also develop custom convex relaxations for the general operations that make up speech preprocessing. Our evaluation across a number of recurrent architectures in computer vision and speech domains shows that these networks are out of reach for existing methods as these are an order of magnitude slower than R2, while R2 successfully verified robustness in many cases.
Enhancing Resilience of Deep Learning Networks by Means of Transferable Adversaries
Seiler, Moritz, Trautmann, Heike, Kerschke, Pascal
Artificial neural networks in general and deep learning networks in particular established themselves as popular and powerful machine learning algorithms. While the often tremendous sizes of these networks are beneficial when solving complex tasks, the tremendous number of parameters also causes such networks to be vulnerable to malicious behavior such as adversarial perturbations. These perturbations can change a model's classification decision. Moreover, while single-step adversaries can easily be transferred from network to network, the transfer of more powerful multi-step adversaries has - usually -- been rather difficult. In this work, we introduce a method for generating strong ad-versaries that can easily (and frequently) be transferred between different models. This method is then used to generate a large set of adversaries, based on which the effects of selected defense methods are experimentally assessed. At last, we introduce a novel, simple, yet effective approach to enhance the resilience of neural networks against adversaries and benchmark it against established defense methods. In contrast to the already existing methods, our proposed defense approach is much more efficient as it only requires a single additional forward-pass to achieve comparable performance results.
CLOCS: Contrastive Learning of Cardiac Signals
Kiyasseh, Dani, Zhu, Tingting, Clifton, David A.
The healthcare industry generates troves of unlabelled physiological data. This data can be exploited via contrastive learning, a self-supervised pre-training method that encourages representations of instances to be similar to one another. We propose a family of contrastive learning methods, CLOCS, that encourages representations across space, time, and patients to be similar to one another. We show that CLOCS consistently outperforms the state-of-the-art methods, BYOL and SimCLR, when performing a linear evaluation of, and fine-tuning on, downstream tasks. We also show that CLOCS achieves strong generalization performance with only 25% of labelled training data. Furthermore, our training procedure naturally generates patient-specific representations that can be used to quantify patient-similarity. At present, the healthcare system is unable to sufficiently leverage the large, unlabelled datasets that it generates on a daily basis. This is partially due to the dependence of deep learning algorithms on high quality labels for good generalization performance. However, arriving at such high quality labels in a clinical setting where physicians are squeezed for time and attention is increasingly difficult. To overcome such an obstacle, self-supervised techniques have emerged as promising methods. These methods exploit the unlabelled dataset to formulate pretext tasks such as predicting the rotation of images (Gidaris et al., 2018), their corresponding colourmap (Larsson et al., 2017), and the arrow of time (Wei et al., 2018). More recently, contrastive learning was introduced as a way to learn representations of instances that share some context. By capturing this high-level shared context (e.g., medical diagnosis), representations become invariant to the differences (e.g., input modalities) between the instances. Contrastive learning can be characterized by three main components: 1) a positive and negative set of examples, 2) a set of transformation operators, and 3) a variant of the noise contrastive estimation loss. Most research in this domain has focused on curating a positive set of examples by exploiting data temporality (Oord et al., 2018), data augmentations (Chen et al., 2020), and multiple views of the same data instance (Tian et al., 2019). These methods are predominantly catered to the image-domain and central to their implementation is the notion that shared context arises from the same instance. We believe this precludes their applicability to the medical domain where physiological time-series are plentiful. Moreover, their interpretation of shared context is limited to data from a common source where that source is the individual data instance.
Precisely Predicting Acute Kidney Injury with Convolutional Neural Network Based on Electronic Health Record Data
Wang, Yu, Bao, JunPeng, Du, JianQiang, Li, YongFeng
The incidence of Acute Kidney Injury (AKI) commonly happens in the Intensive Care Unit (ICU) patients, especially in the adults, which is an independent risk factor affecting short-term and long-term mortality. Though researchers in recent years highlight the early prediction of AKI, the performance of existing models are not precise enough. The objective of this research is to precisely predict AKI by means of Convolutional Neural Network on Electronic Health Record (EHR) data. The data sets used in this research are two public Electronic Health Record (EHR) databases: MIMIC-III and eICU database. In this study, we take several Convolutional Neural Network models to train and test our AKI predictor, which can precisely predict whether a certain patient will suffer from AKI after admission in ICU according to the last measurements of the 16 blood gas and demographic features. The research is based on Kidney Disease Improving Global Outcomes (KDIGO) criteria for AKI definition. Our work greatly improves the AKI prediction precision, and the best AUROC is up to 0.988 on MIMIC-III data set and 0.936 on eICU data set, both of which outperform the state-of-art predictors. And the dimension of the input vector used in this predictor is much fewer than that used in other existing researches. Compared with the existing AKI predictors, the predictor in this work greatly improves the precision of early prediction of AKI by using the Convolutional Neural Network architecture and a more concise input vector. Early and precise prediction of AKI will bring much benefit to the decision of treatment, so it is believed that our work is a very helpful clinical application.
Bayesian Generative Models for Knowledge Transfer in MRI Semantic Segmentation Problems
Kuzina, Anna, Egorov, Evgenii, Burnaev, Evgeny
Automatic segmentation methods based on deep learning have recently demonstrated state-of-the-art performance, outperforming the ordinary methods. Nevertheless, these methods are inapplicable for small datasets, which are very common in medical problems. To this end, we propose a knowledge transfer method between diseases via the Generative Bayesian Prior network. Our approach is compared to a pre-train approach and random initialization and obtains the best results in terms of Dice Similarity Coefficient metric for the small subsets of the Brain Tumor Segmentation 2018 database (BRATS2018).
COVID-19 growth prediction using multivariate long short term memory
Coronavirus disease (COVID-19) spread forecasting is an important task to track the growth of the pandemic. Existing predictions are merely based on qualitative analyses and mathematical modeling. The use of available big data with machine learning is still limited in COVID-19 growth prediction even though the availability of data is abundance. To make use of big data in the prediction using deep learning, we use long short-term memory (LSTM) method to learn the correlation of COVID-19 growth over time. The structure of an LSTM layer is searched heuristically until the best validation score is achieved. First, we trained training data containing confirmed cases from around the globe. We achieved favorable performance compared with that of the recurrent neural network (RNN) method with a comparable low validation error. The evaluation is conducted based on graph visualization and root mean squared error (RMSE). We found that it is not easy to achieve the same quantity of confirmed cases over time. However, LSTM provide a similar pattern between the actual cases and prediction. In the future, our proposed prediction can be used for anticipating forthcoming pandemics. The code is provided here: https://github.com/cbasemaster/lstmcorona
Live Trojan Attacks on Deep Neural Networks
Costales, Robby, Mao, Chengzhi, Norwitz, Raphael, Kim, Bryan, Yang, Junfeng
Like all software systems, the execution of deep learning models is dictated in part by logic represented as data in memory. For decades, attackers have exploited traditional software programs by manipulating this data. We propose a live attack on deep learning systems that patches model parameters in memory to achieve predefined malicious behavior on a certain set of inputs. By minimizing the size and number of these patches, the attacker can reduce the amount of network communication and memory overwrites, with minimal risk of system malfunctions or other detectable side effects. We demonstrate the feasibility of this attack by computing efficient patches on multiple deep learning models. We show that the desired trojan behavior can be induced with a few small patches and with limited access to training data. We describe the details of how this attack is carried out on real systems and provide sample code for patching TensorFlow model parameters in Windows and in Linux. Lastly, we present a technique for effectively manipulating entropy on perturbed inputs to bypass STRIP, a state-of-the-art run-time trojan detection technique.
Self-Supervised Representation Learning on Document Images
Cosma, Adrian, Ghidoveanu, Mihai, Panaitescu-Liess, Michael, Popescu, Marius
While previous approaches explore the effect of self-supervision on natural images, we show that patch-based pre-training performs poorly on document images because of their different structural properties and poor intra-sample semantic information. We propose two context-aware alternatives to improve performance on the Tobacco-3482 image classification task. We also propose a novel method for self-supervision, which makes use of the inherent multi-modality of documents (image and text), which performs better than other popular self-supervised methods, including supervised ImageNet pre-training, on document image classification scenarios with a limited amount of data.