Deep Learning
Sequential Random Network for Fine-grained Image Classification
Li, Chaorong, Zhang, Malu, Huang, Wei, Qin, Fengqing, Zeng, Anping, Huang, Yuanyuan
Deep Convolutional Neural Network (DCNN) and Transformer have achieved remarkable successes in image recognition. However, their performance in fine-grained image recognition is still difficult to meet the requirements of actual needs. This paper proposes a Sequence Random Network (SRN) to enhance the performance of DCNN. The output of DCNN is one-dimensional features. This one-dimensional feature abstractly represents image information, but it does not express well the detailed information of image. To address this issue, we use the proposed SRN which composed of BiLSTM and several Tanh-Dropout blocks (called BiLSTM-TDN), to further process DCNN one-dimensional features for highlighting the detail information of image. After the feature transform by BiLSTM-TDN, the recognition performance has been greatly improved. We conducted the experiments on six fine-grained image datasets. Except for FGVC-Aircraft, the accuracies of the proposed methods on the other datasets exceeded 99%. Experimental results show that BiLSTM-TDN is far superior to the existing state-of-the-art methods. In addition to DCNN, BiLSTM-TDN can also be extended to other models, such as Transformer.
BDD4BNN: A BDD-based Quantitative Analysis Framework for Binarized Neural Networks
Zhang, Yedi, Zhao, Zhe, Chen, Guangke, Song, Fu, Chen, Taolue
Verifying and explaining the behavior of neural networks is becoming increasingly important, especially when they are deployed in safety-critical applications. In this paper, we study verification problems for Binarized Neural Networks (BNNs), the 1-bit quantization of general real-numbered neural networks. Our approach is to encode BNNs into Binary Decision Diagrams (BDDs), which is done by exploiting the internal structure of the BNNs. In particular, we translate the input-output relation of blocks in BNNs to cardinality constraints which are then encoded by BDDs. Based on the encoding, we develop a quantitative verification framework for BNNs where precise and comprehensive analysis of BNNs can be performed. We demonstrate the application of our framework by providing quantitative robustness analysis and interpretability for BNNs. We implement a prototype tool BDD4BNN and carry out extensive experiments which confirm the effectiveness and efficiency of our approach.
Auction Based Clustered Federated Learning in Mobile Edge Computing System
Lu, Renhao, Zhang, Weizhe, Li, Qiong, Zhong, Xiaoxiong, Vasilakos, Athanasios V.
In recent years, mobile clients' computing ability and storage capacity have greatly improved, efficiently dealing with some applications locally. Federated learning is a promising distributed machine learning solution that uses local computing and local data to train the Artificial Intelligence (AI) model. Combining local computing and federated learning can train a powerful AI model under the premise of ensuring local data privacy while making full use of mobile clients' resources. However, the heterogeneity of local data, that is, Non-independent and identical distribution (Non-IID) and imbalance of local data size, may bring a bottleneck hindering the application of federated learning in mobile edge computing (MEC) system. Inspired by this, we propose a cluster-based clients selection method that can generate a federated virtual dataset that satisfies the global distribution to offset the impact of data heterogeneity and proved that the proposed scheme could converge to an approximate optimal solution. Based on the clustering method, we propose an auction-based clients selection scheme within each cluster that fully considers the system's energy heterogeneity and gives the Nash equilibrium solution of the proposed scheme for balance the energy consumption and improving the convergence rate. The simulation results show that our proposed selection methods and auction-based federated learning can achieve better performance with the Convolutional Neural Network model (CNN) under different data distributions.
Explaining Network Intrusion Detection System Using Explainable AI Framework
Cybersecurity is a domain where the data distribution is constantly changing with attackers exploring newer patterns to attack cyber infrastructure. Intrusion detection system is one of the important layers in cyber safety in today's world. Machine learning based network intrusion detection systems started showing effective results in recent years. With deep learning models, detection rates of network intrusion detection system are improved. More accurate the model, more the complexity and hence less the interpretability. Deep neural networks are complex and hard to interpret which makes difficult to use them in production as reasons behind their decisions are unknown. In this paper, we have used deep neural network for network intrusion detection and also proposed explainable AI framework to add transparency at every stage of machine learning pipeline. This is done by leveraging Explainable AI algorithms which focus on making ML models less of black boxes by providing explanations as to why a prediction is made. Explanations give us measurable factors as to what features influence the prediction of a cyberattack and to what degree. These explanations are generated from SHAP, LIME, Contrastive Explanations Method, ProtoDash and Boolean Decision Rules via Column Generation. We apply these approaches to NSL KDD dataset for intrusion detection system and demonstrate results.
The Achilles' heel of AI might be its big carbon footprint
A few months ago, Generative Pre-Trained Transformer-3, or GPT-3, the biggest artificial intelligence (AI) model in history and the most powerful language model ever, was launched with much fanfare by OpenAI, a San Francisco-based AI lab. Over the last few years, one of the biggest trends in natural language processing (NLP) has been the increasing size of language models (LMs), as measured by the size of training data and the number of parameters. The 2018-released BERT, which was then considered the best-in-class NLP model, was trained on a dataset of 3 billion words. The XLNet model that outperformed BERT was based on a training set of 32 billion words. Shortly thereafter, GPT-2 was trained on a dataset of 40 billion words. Dwarfing all these, GPT-3 was trained on a weighted dataset of roughly 500 billion words.
Using the power of artificial intelligence to detect disease
A large international collaboration, led by A/Prof Xiu Ying Wang and Prof Manuel Graeber of the University of Sydney, has developed an innovative, advanced artificial intelligence (AI) application, PathoFusion, that could be used for the examination of routine tissue samples in order to identify indications of cancer. The research melds contributions from computer scientists, neuropathologists, neuosurgeons, medical oncologists and medical imaging scientists. ANSTO's Prof Richard Banati, a Professor of Medical Radiation Sciences/Medical Imaging, who studies the brain's innate immune system using advanced medical imaging techniques, is a co-author on the paper published in the journal, Cancers. "The idea behind PathoFusion was to create a novel advanced deep learning model to recognize malignant features and immune response markers, independent of human intervention, and map them simultaneously in a digital image," explained Banati. Scientists specifically designed a bifocal deep learning framework which is analogous to how a microscopist works in histopathology image analysis.
Intel hooks up with Deci for deep learning
As one of the first companies to participate in Intel Ignite startup accelerator, Deci will now work with Intel to deploy innovative AI technologies to mutual customers. The collaboration takes helps enable deep learning inference at scale on Intel CPUs, reducing costs and latency, and enabling new applications of deep learning inference. New deep learning tasks can be performed in a real-time environment on edge devices and companies that use large scale inference scenarios can dramatically cut cloud or datacenter cost, simply by changing the inference hardware from GPU to Intel CPU. "By optimizing the AI models that run on Intel's hardware, Deci enables customers to get even more speed and will allow for cost-effective and more general deep learning use cases on Intel CPUs," says Deci CEO and co-founder Yonatan Geifman. Deci and Intel's collaboration began with MLPerf where on several Intel CPUs, Deci's AutoNAC (Automated Neural Architecture Construction) technology accelerated the inference speed of the well-known ResNet-50 neural network, reducing the submitted models' latency by a factor of up to 11.8x and increasing throughput by up to 11x.
Prediction Models for AKI in ICU: A Comparative Study
Purpose: To assess the performance of models for early prediction of acute kidney injury (AKI) in the Intensive Care Unit (ICU) setting. Patients and Methods: Data were collected from the Medical Information Mart for Intensive Care (MIMIC)-III database for all patients aged 18 years who had their serum creatinine (SCr) level measured for 72 h following ICU admission. Those with existing conditions of kidney disease upon ICU admission were excluded from our analyses. Seventeen predictor variables comprising patient demographics and physiological indicators were selected on the basis of the Kidney Disease Improving Global Outcomes (KDIGO) and medical literature. Six models from three types of methods were tested: Logistic Regression (LR), Support Vector Machines (SVM), Random Forest (RF), eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Decision Machine (LightGBM), and Convolutional Neural Network (CNN).
Step-By-Step Implementation of GANs on Custom Image Data in PyTorch: Part 2
In Part 1 on GANs, we started to build intuition regarding what GANs are, why we need them, and how the entire point behind training GANs is to create a generator model that knows how to convert a random noise vector into a (beautiful) almost real image. Since we have already discussed the pseudocode in great depth in Part 1, be sure to check that out as there will be a lot of references to it! In case you would like to follow along, here is the Github Notebook containing the source code for training GANs using the PyTorch framework. The whole idea behind training a GAN network is to obtain a Generator network (with most optimal model weights and layers, etc.) that is excellent at spewing out fakes that look like real! Note: I would like to take a moment to truly appreciate Nathan Inkawhich for writing a superb article explaining the inner workings of DCGANs and the official Github repository for Pytorch that helped me with the code implementations, especially the network architectures for both Generator and Discriminator. Hopefully, the explanations I have presented in this article help you gain even further clarity (than already present in the aforementioned blogs) regarding GANs and implement them even better for your own use-case! If this in-depth educational content is useful for you, you can subscribe to our AI research mailing list to be alerted when we release new material.
Going Deep on Deep Learning
In a field with a tendency to jump from buzzword to buzzword, deep learning has proven to be a consistent driver of discovery and scientific progress -- but also a frequent source of confusion, and even trepidation. Well, TDS is here to help. A good place to start? This comprehensive introduction to Bayesian deep learning by Joris Baan, which he wrote with the explicit goal of bridging the gap between basic probability theory and cutting-edge research. You'd be hard-pressed to find a better article to read next than Lina Faik's patient explanation of deep Q learning, reinforcement learning, and how the two are powering new and exciting real-world applications.