Goto

Collaborating Authors

 Deep Learning


FedGEMS: Federated Learning of Larger Server Models via Selective Knowledge Fusion

arXiv.org Artificial Intelligence

Today data is often scattered among billions of resource-constrained edge devices with security and privacy constraints. Federated Learning (FL) has emerged as a viable solution to learn a global model while keeping data private, but the model complexity of FL is impeded by the computation resources of edge nodes. In this work, we investigate a novel paradigm to take advantage of a powerful server model to break through model capacity in FL. By selectively learning from multiple teacher clients and itself, a server model develops in-depth knowledge and transfers its knowledge back to clients in return to boost their respective performance. Our proposed framework achieves superior performance on both server and client models and provides several advantages in a unified framework, including flexibility for heterogeneous client architectures, robustness to poisoning attacks, and communication efficiency between clients and server. By bridging FL effectively with larger server model training, our proposed paradigm paves ways for robust and continual knowledge accumulation from distributed and private data. Nowadays, large models with tremendous parameters trained with sufficient computation power are indispensable in Artificial Intelligence (AI), such as AlphaGo (Silver et al., 2016), Alphafold (Senior et al., 2020) and GPT -3 (Brown et al., 2020).


A channel attention based MLP-Mixer network for motor imagery decoding with EEG

arXiv.org Artificial Intelligence

Convolutional neural networks (CNNs) and their variants have been successfully applied to the electroencephalogram (EEG) based motor imagery (MI) decoding task. However, these CNN-based algorithms generally have limitations in perceiving global temporal dependencies of EEG signals. Besides, they also ignore the diverse contributions of different EEG channels to the classification task. To address such issues, a novel channel attention based MLP-Mixer network (CAMLP-Net) is proposed for EEG-based MI decoding. Specifically, the MLP-based architecture is applied in this network to capture the temporal and spatial information. The attention mechanism is further embedded into MLP-Mixer to adaptively exploit the importance of different EEG channels. Therefore, the proposed CAMLP-Net can effectively learn more global temporal and spatial information. The experimental results on the newly built MI-2 dataset indicate that our proposed CAMLP-Net achieves superior classification performance over all the compared algorithms.


Can Q-learning solve Multi Armed Bantids?

arXiv.org Artificial Intelligence

When a reinforcement learning (RL) method has to decide between several optional policies by solely looking at the received reward, it has to implicitly optimize a Multi-Armed-Bandit (MAB) problem. This arises the question: are current RL algorithms capable of solving MAB problems? We claim that the surprising answer is no. In our experiments we show that in some situations they fail to solve a basic MAB problem, and in many common situations they have a hard time: They suffer from regression in results during training, sensitivity to initialization and high sample complexity. We claim that this stems from variance differences between policies, which causes two problems: The first problem is the "Boring Policy Trap" where each policy have a different implicit exploration depends on its rewards variance, and leaving a boring, or low variance, policy is less likely due to its low implicit exploration. The second problem is the "Manipulative Consultant" problem, where value-estimation functions used in deep RL algorithms such as DQN or deep Actor Critic methods, maximize estimation precision rather than mean rewards, and have a better loss in low-variance policies, which cause the network to converge to a sub-optimal policy. Cognitive experiments on humans showed that noised reward signals may paradoxically improve performance. We explain this using the aforementioned problems, claiming that both humans and algorithms may share similar challenges in decision making. Inspired by this result, we propose the Adaptive Symmetric Reward Noising (ASRN) method, by which we mean equalizing the rewards variance across different policies, thus avoiding the two problems without affecting the environment's mean rewards behavior. We demonstrate that the ASRN scheme can dramatically improve the results.


Dual-branch Attention-In-Attention Transformer for single-channel speech enhancement

arXiv.org Artificial Intelligence

Curriculum learning begins to thrive in the speech enhancement area, which decouples the original spectrum estimation task into multiple easier sub-tasks to achieve better performance. Motivated by that, we propose a dual-branch attention-in-attention transformer dubbed DB-AIAT to handle both coarse- and fine-grained regions of the spectrum in parallel. From a complementary perspective, a magnitude masking branch is proposed to coarsely estimate the overall magnitude spectrum, and simultaneously a complex refining branch is elaborately designed to compensate for the missing spectral details and implicitly derive phase information. Within each branch, we propose a novel attention-in-attention transformer-based module to replace the conventional RNNs and temporal convolutional networks for temporal sequence modeling. Specifically, the proposed attention-in-attention transformer consists of adaptive temporal-frequency attention transformer blocks and an adaptive hierarchical attention module, aiming to capture long-term temporal-frequency dependencies and further aggregate global hierarchical contextual information. Experimental results on Voice Bank + DEMAND demonstrate that DB-AIAT yields state-of-the-art performance (e.g., 3.31 PESQ, 95.6% STOI and 10.79dB SSNR) over previous advanced systems with a relatively small model size (2.81M).


A Real-Time Energy and Cost Efficient Vehicle Route Assignment Neural Recommender System

arXiv.org Machine Learning

This paper presents a neural network recommender system algorithm for assigning vehicles to routes based on energy and cost criteria. In this work, we applied this new approach to efficiently identify the most cost-effective medium and heavy duty truck (MDHDT) powertrain technology, from a total cost of ownership (TCO) perspective, for given trips. We employ a machine learning based approach to efficiently estimate the energy consumption of various candidate vehicles over given routes, defined as sequences of links (road segments), with little information known about internal dynamics, i.e using high level macroscopic route information. A complete recommendation logic is then developed to allow for real-time optimum assignment for each route, subject to the operational constraints of the fleet. We show how this framework can be used to (1) efficiently provide a single trip recommendation with a top-$k$ vehicles star ranking system, and (2) engage in more general assignment problems where $n$ vehicles need to be deployed over $m \leq n$ trips. This new assignment system has been deployed and integrated into the POLARIS Transportation System Simulation Tool for use in research conducted by the Department of Energy's Systems and Modeling for Accelerated Research in Transportation (SMART) Mobility Consortium


Exploring Architectural Ingredients of Adversarially Robust Deep Neural Networks

arXiv.org Machine Learning

Deep neural networks (DNNs) are known to be vulnerable to adversarial attacks. A range of defense methods have been proposed to train adversarially robust DNNs, among which adversarial training has demonstrated promising results. However, despite preliminary understandings developed for adversarial training, it is still not clear, from the architectural perspective, what configurations can lead to more robust DNNs. In this paper, we address this gap via a comprehensive investigation on the impact of network width and depth on the robustness of adversarially trained DNNs. Specifically, we make the following key observations: 1) more parameters (higher model capacity) does not necessarily help adversarial robustness; 2) reducing capacity at the last stage (the last group of blocks) of the network can actually improve adversarial robustness; and 3) under the same parameter budget, there exists an optimal architectural configuration for adversarial robustness. We also provide a theoretical analysis explaning why such network configuration can help robustness. These architectural insights can help design adversarially robust DNNs.


@Radiology_AI

#artificialintelligence

To evaluate two settings (noise reduction of 50% or 75%) of a deep learning (DL) reconstruction model relative to each other and to conventional MR image reconstructions on clinical orthopedic MRI datasets. This retrospective study included 54 patients who underwent two-dimensional fast spin-echo MRI for hip (n 22; mean age, 44 years 13 [standard deviation]; nine men) or shoulder (n 32; mean age, 56 years 17; 17 men) conditions between March 2019 and June 2020. MR images were reconstructed with conventional methods and the vendor-provided and commercially available DL model applied with 50% and 75% noise reduction settings (DL 50 and DL 75, respectively). Quantitative analytics, including relative anatomic edge sharpness, relative signal-to-noise ratio (rSNR), and relative contrast-to-noise ratio (rCNR) were computed for each dataset. In addition, the image sets were randomized, blinded, and presented to three board-certified musculoskeletal radiologists for ranking based on overall image quality and diagnostic confidence.


Top AI Software

#artificialintelligence

Clearly, today's best artificial intelligence software is driving change: hardly a day goes by without artificial intelligence (AI) software introducing new and improved capabilities. As features appear and use cases expand, organizations turn to AI applications to gain competitive advantage. AI software capabilities typically fall into several core areas: machine learning (ML), deep learning, predictive analytics, machine vision, robotic process automation (RPA), smart assistants and chatbots. With the current focus on digital transformation, systems are changing everything from business forecasting and supply chain automation to marketing/sales and customer support. They're ushering in smarter business and IT frameworks that can act and react to events in more agile and flexible ways.


The Warmup Guide to Hugging Face

#artificialintelligence

Since it was founded, the startup, Hugging Face, has created several open-source libraries for NLP-based tokenizers and transformers. One of their libraries, the Hugging Face transformers package, is an immensely popular Python library providing over 32 pre-trained models that are extraordinarily useful for a variety of natural language processing (NLP) tasks. It was created to enable general-purpose architectures such as BERT, GPT-2, XLNet, XLM, DistilBERT, and RoBERTa for Natural Language Understanding (NLU) and Natural Language Generation (NLG) to perform tasks like text classification, information extraction, and text generation. Tasks performed by this library include classification, information extraction, question answering, summarization, translation, and text generation in over 100 languages. Its ultimate goal is to make cutting-edge NLP easier to use for everyone.


Scientists show how AI may spot unseen signs of heart failure

#artificialintelligence

A special artificial intelligence (AI)-based computer algorithm created by Mount Sinai researchers was able to learn how to identify subtle changes in electrocardiograms (also known as ECGs or EKGs) to predict whether a patient was experiencing heart failure. "We showed that deep-learning algorithms can recognize blood pumping problems on both sides of the heart from ECG waveform data," said Benjamin S. Glicksberg, Ph.D., Assistant Professor of Genetics and Genomic Sciences, a member of the Hasso Plattner Institute for Digital Health at Mount Sinai, and a senior author of the study published in the Journal of the American College of Cardiology: Cardiovascular Imaging. "Ordinarily, diagnosing these type of heart conditions requires expensive and time-consuming procedures. We hope that this algorithm will enable quicker diagnosis of heart failure." The study was led by Akhil Vaid, MD, a postdoctoral scholar who works in both the Glicksberg lab and one led by Girish N. Nadkarni, MD, MPH, CPH, Associate Professor of Medicine at the Icahn School of Medicine at Mount Sinai, Chief of the Division of Data-Driven and Digital Medicine (D3M), and a senior author of the study.