Goto

Collaborating Authors

 Deep Learning


Towards Class Imbalance in Federated Learning

arXiv.org Machine Learning

Federated learning (FL) is a promising approach for training decentralized data located on local client devices while improving efficiency and privacy. However, the distribution and quantity of the training data on the clients' side may lead to significant challenges such as data imbalance and non-IID (non-independent and identically distributed) data, which could greatly impact the performance of the common model. While much effort has been devoted to helping FL models converge when encountering non-IID data, the imbalance issue has not been sufficiently addressed. In particular, as FL training is executed by exchanging gradients in an encrypted form, the training data is not completely observable to either clients or server, and previous methods for data imbalance do not perform well for FL. Therefore, it is crucial to design new methods for detecting data imbalance in FL and mitigating its impact. In this work, we propose a monitoring scheme that can infer the composition proportion of training data for each FL round, and design a new loss function -- Ratio Loss to mitigate the impact of the imbalance. Our experiments demonstrate the importance of detecting data imbalance and taking measures as early as possible in FL training, and the effectiveness of our method in mitigating the impact. Our method is shown to significantly outperform previous methods, while maintaining client privacy.


Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text Classification

arXiv.org Machine Learning

Extreme multi-label text classification (XMTC) is a task for tagging a given text with the most relevant labels from an extremely large label set. We propose a novel deep learning method called APLC-XLNet. Our approach fine-tunes the recently released generalized autoregressive pretrained model (XLNet) to learn a dense representation for the input text. We propose Adaptive Probabilistic Label Clusters (APLC) to approximate the cross entropy loss by exploiting the unbalanced label distribution to form clusters that explicitly reduce the computational time. Our experiments, carried out on five benchmark datasets, show that our approach has achieved new state-of-the-art results on four benchmark datasets. Our source code is available publicly at https://github.com/huiyegit/APLC_XLNet.


Decentralized Reinforcement Learning: Global Decision-Making via Local Economic Transactions

arXiv.org Machine Learning

This paper seeks to establish a framework for directing a society of simple, specialized, self-interested agents to solve what traditionally are posed as monolithic single-agent sequential decision problems. What makes it challenging to use a decentralized approach to collectively optimize a central objective is the difficulty in characterizing the equilibrium strategy profile of non-cooperative games. To overcome this challenge, we design a mechanism for defining the learning environment of each agent for which we know that the optimal solution for the global objective coincides with a Nash equilibrium strategy profile of the agents optimizing their own local objectives. The society functions as an economy of agents that learn the credit assignment process itself by buying and selling to each other the right to operate on the environment state. We derive a class of decentralized reinforcement learning algorithms that are broadly applicable not only to standard reinforcement learning but also for selecting options in semi-MDPs and dynamically composing computation graphs. Lastly, we demonstrate the potential advantages of a society's inherent modular structure for more efficient transfer learning.


Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors

arXiv.org Machine Learning

Bayesian neural networks (BNNs) demonstrate promising success in improving the robustness and uncertainty quantification of modern deep learning. However, they generally struggle with underfitting at scale and parameter efficiency. On the other hand, deep ensembles have emerged as alternatives for uncertainty quantification that, while outperforming BNNs on certain problems, also suffer from efficiency issues. It remains unclear how to combine the strengths of these two approaches and remediate their common issues. To tackle this challenge, we propose a rank-1 parameterization of BNNs, where each weight matrix involves only a distribution on a rank-1 subspace. We also revisit the use of mixture approximate posteriors to capture multiple modes, where unlike typical mixtures, this approach admits a significantly smaller memory increase (e.g., only a 0.4% increase for a ResNet-50 mixture of size 10). We perform a systematic empirical study on the choices of prior, variational posterior, and methods to improve training. For ResNet-50 on ImageNet, Wide ResNet 28-10 on CIFAR-10/100, and an RNN on MIMIC-III, rank-1 BNNs achieve state-of-the-art performance across log-likelihood, accuracy, and calibration on the test sets and out-of-distribution variants.


How to Train Your Energy-Based Model for Regression

arXiv.org Machine Learning

Energy-based models (EBMs) have become increasingly popular within computer vision in recent years. While they are commonly employed for generative image modeling, recent work has applied EBMs also for regression tasks, achieving state-of-the-art performance on object detection and visual tracking. Training EBMs is however known to be challenging. While a variety of different techniques have been explored for generative modeling, the application of EBMs to regression is not a well-studied problem. How EBMs should be trained for best possible regression performance is thus currently unclear. We therefore accept the task of providing the first detailed study of this problem. To that end, we propose a simple yet highly effective extension of noise contrastive estimation, and carefully compare its performance to six popular methods from literature on the tasks of 1D regression and object detection. The results of this comparison suggest that our training method should be considered the go-to approach. We also apply our method to the visual tracking task, achieving state-of-the-art performance on five datasets. Notably, our tracker achieves 63.7% AUC on LaSOT and 78.7% Success on TrackingNet. Code is available at https://github.com/fregu856/ebms_regression.


Why GPT-3 Heralds a Democratic Revolution in Tech

#artificialintelligence

Ph.D. in mathematics, named on a Forbes 30 under 30 list. He runs Contentyze, a content generation platform, and is passionate about educational content. Ph.D. in mathematics, named on a Forbes 30 under 30 list. He runs Contentyze, a content generation platform, and is passionate about educational content. GPT-3, a machine learning model from OpenAI, has taken the world by storm over the last couple of weeks.


OpenAI GPT-3: How It Works and Why It Matters - DZone AI

#artificialintelligence

You have probably heard about an innovative language model called GPT3. The hype is so overwhelming that we decided to research its core and the consequences for the tech players. Let's explore whether the language deserves this much attention and what makes it so exceptional. GPT-3 is a text generating neural network that was released in June 2020 and tested for $14 million. Its creator is the AI research agency OpenAI headed by Sam Altman, Marc Benioff, Elon Musk, and Reid Hoffman. The language is based on 175 million parameters and is by far more accurate than its predecessors.


Sentiment Analysis using Deep Learning

#artificialintelligence

The growth of the internet due to social networks such as facebook, twitter, Linkedin, instagram etc. has led to significant users interaction and has empowered users to express their opinions about products, services, events, their preferences among others. It has also provided opportunities to the users to share their wisdom and experiences with each other. The faster development of social networks is causing explosive growth of digital content. It has turned online opinions, blogs, tweets, and posts into a very valuable asset for the corporates to get insights from the data and plan their strategy. Business organizations need to process and study these sentiments to investigate data and to gain business insights(Yadav & Vishwakarma, 2020).


Protecting personal data in AI – new ICO guidance

#artificialintelligence

Developers of artificial intelligence ("AI") continue to push the technology to the next level; not a week goes by without a news story claiming that the next big thing has arrived. One popular recent example is OpenAI's latest machine learning system called GPT-3, which was trained on 45TB of text data and has been pounced upon by fellow developers and commentators alike. Its potential is vast but we are still in the early days of the AI revolution. However, as the AI available to be used grows more powerful, it becomes ever more important to consider the ethical and legal issues involved. In this context, the Information Commissioner's Office ("ICO"), the independent authority in the UK for enforcing data protection laws, has released its "Guidance on AI and Data Protection".


Why You Should Start Your Deep Learning Journey With PyTorch

#artificialintelligence

There's no denying the fact that Deep Learning as we know it, how awesome it is when we can see that with minimal or no human-intervention a job can be done. Since, Machine Learning, Deep Learning is dubbed to be one of the sexiest jobs of the 21st century(hyped?) so there has to be some starting point, a sort of a roadmap that you can follow to reach to the other side. Luckily, we can now approach it relatively easier with modern frameworks like Tensorflow, PyTorch which gives you a high-level interface to build awesome stuff! Let's discuss why you should start with PyTorch. That means line-by-line execution of the code and simultaneous building of the computation graphs just like in python.