Education
Defense Through Diverse Directions
Bender, Christopher M., Li, Yang, Shi, Yifeng, Reiter, Michael K., Oliva, Junier B.
In this work we develop a novel Bayesian neural network methodology to achieve strong adversarial robustness without the need for online adversarial training. Unlike previous efforts in this direction, we do not rely solely on the stochasticity of network weights by minimizing the divergence between the learned parameter distribution and a prior. Instead, we additionally require that the model maintain some expected uncertainty with respect to all input covariates. We demonstrate that by encouraging the network to distribute evenly across inputs, the network becomes less susceptible to localized, brittle features which imparts a natural robustness to targeted perturbations. We show empirical robustness on several benchmark datasets.
Meta Pseudo Labels
Pham, Hieu, Xie, Qizhe, Dai, Zihang, Le, Quoc V.
Many training algorithms of a deep neural network can be interpreted as minimizing the cross entropy loss between the prediction made by the network and a target distribution. In supervised learning, this target distribution is typically the ground-truth one-hot vector. In semi-supervised learning, this target distribution is typically generated by a pre-trained teacher model to train the main network. In this work, instead of using such predefined target distributions, we show that learning to adjust the target distribution based on the learning state of the main network can lead to better performances. In particular, we propose an efficient meta-learning algorithm, which encourages the teacher to adjust the target distributions of training examples in the manner that improves the learning of the main network. The teacher is updated by policy gradients computed by evaluating the main network on a held-out validation set. Our experiments demonstrate substantial improvements over strong baselines and establish state-ofthe-art performance on CIFAR-10, SVHN, and ImageNet. For instance, with ResNets on small datasets, we achieve 96.1% on CIFAR-10 with 4,000 labeled examples and 73.9% top-1 on ImageNet with 10% examples. Meanwhile, with EfficientNet on full datasets plus extra unlabeled data, we attain 98.6% accuracy on CIFAR-10 and 86.9% top-1 accuracy on ImageNet.
Neural Networks and Polynomial Regression. Demystifying the Overparametrization Phenomena
Emschwiller, Matt, Gamarnik, David, Kฤฑzฤฑldaฤ, Eren C., Zadik, Ilias
In the context of neural network models, overparametrization refers to the phenomena whereby these models appear to generalize well on the unseen data, even though the number of parameters significantly exceeds the sample sizes, and the model perfectly fits the in-training data. A conventional explanation of this phenomena is based on self-regularization properties of algorithms used to train the data. In this paper we prove a series of results which provide a somewhat diverging explanation. Adopting a teacher/student model where the teacher network is used to generate the predictions and student network is trained on the observed labeled data, and then tested on out-of-sample data, we show that any student network interpolating the data generated by a teacher network generalizes well, provided that the sample size is at least an explicit quantity controlled by data dimension and approximation guarantee alone, regardless of the number of internal nodes of either teacher or student network. Our claim is based on approximating both teacher and student networks by polynomial (tensor) regression models with degree depending on the desired accuracy and network depth only. Such a parametrization notably does not depend on the number of internal nodes. Thus a message implied by our results is that parametrizing wide neural networks by the number of hidden nodes is misleading, and a more fitting measure of parametrization complexity is the number of regression coefficients associated with tensorized data. In particular, this somewhat reconciles the generalization ability of neural networks with more classical statistical notions of data complexity and generalization bounds. Our empirical results on MNIST and Fashion-MNIST datasets indeed confirm that tensorized regression achieves a good out-of-sample performance, even when the degree of the tensor is at most two.
Incorporating Relational Background Knowledge into Reinforcement Learning via Differentiable Inductive Logic Programming
Relational Reinforcement Learning (RRL) can offers various desirable features. Most importantly, it allows for incorporating expert knowledge into the learning, and hence leading to much faster learning and better generalization compared to the standard deep reinforcement learning. However, most of the existing RRL approaches are either incapable of incorporating expert background knowledge (e.g., in the form of explicit predicate language) or are not able to learn directly from non-relational data such as image. In this paper, we propose a novel deep RRL based on a differentiable Inductive Logic Programming (ILP) that can effectively learn relational information from image and present the state of the environment as first order logic predicates. Additionally, it can take the expert background knowledge and incorporate it into the learning problem using appropriate predicates. The differentiable ILP allows an end to end optimization of the entire framework for learning the policy in RRL. We show the efficacy of this novel RRL framework using environments such as BoxWorld, GridWorld as well as relational reasoning for the Sort-of-CLEVR dataset.
Anticipatory Psychological Models for Quickest Change Detection: Human Sensor Interaction
We consider anticipatory psychological models for human decision makers and their effect on sequential decision making. From a decision theoretic point of view, such models are time inconsistent meaning that Bellman's principle of optimality does not hold. The aim of this paper is to study how such an anxiety-based anticipatory utility can affect sequential decision making, such as quickest change detection, in multi-agent systems. We show that the interaction between anticipation-driven agents and sequential decision maker results in unusual (nonconvex) structure of the optimal decision policy. The methodology yields a useful mathematical framework for sensor interaction involving a human decision maker (with behavioral economics constraints) and a sensor equipped with automated sequential detector.
Evolutionary Population Curriculum for Scaling Multi-Agent Reinforcement Learning
Long, Qian, Zhou, Zihan, Gupta, Abhibav, Fang, Fei, Wu, Yi, Wang, Xiaolong
In multi-agent games, the complexity of the environment can grow exponentially as the number of agents increases, so it is particularly challenging to learn good policies when the agent population is large. In this paper, we introduce Evolutionary Population Curriculum (EPC), a curriculum learning paradigm that scales up Multi-Agent Reinforcement Learning (MARL) by progressively increasing the population of training agents in a stage-wise manner. Furthermore, EPC uses an evolutionary approach to fix an objective misalignment issue throughout the curriculum: agents successfully trained in an early stage with a small population are not necessarily the best candidates for adapting to later stages with scaled populations. Concretely, EPC maintains multiple sets of agents in each stage, performs mix-and-match and fine-tuning over these sets and promotes the sets of agents with the best adaptability to the next stage. We implement EPC on a popular MARL algorithm, MADDPG, and empirically show that our approach consistently outperforms baselines by a large margin as the number of agents grows exponentially. The project page is https://sites.google.com/view/epciclr2020.
How High-Performing Companies Develop and Scale AI
In the latest McKinsey Global Survey on AI we noted a significant year-over-year jump in companies using AI across multiple areas of the business. And while most survey respondents said their companies have gained value from AI, some are attaining greater scale, revenue increases, and cost savings than the rest. Based on our research and experience, this is no accident; how companies build their business strategy, what foundations they put in place, and how they tackle AI adoption in the workplace can all impact their potential for transformation. Many companies that have spent years developing AI technologies are facing the stark reality that successfully scaling AI requires more than just deploying AI technology. We find that those companies finding more success in scaling efforts are more likely than others to apply a core set of practices.
EshbanTheLearner/thepersonalmsds
Caution: This timeline is tailored for @EshbanTheLearner and might not be suitable for everyone. Today's Progress: Today I continued with the Google Cloud Platform Big Data and Machine Learning Fundamentals course of Data Engineering with Google Cloud Professional Certificate on Coursera. Today's Progress: Today I enrolled in the Google Cloud Platform Big Data and Machine Learning Fundamentals course of Data Engineering with Google Cloud Professional Certificate on Coursera. Today's Progress: Today I ended and reviewd the Time Series Analysis in Python 2020 course on udemy. Today's Progress: Today I continued the Time Series Analysis in Python 2020 course on udemy.
How Can We Use AI and Chatbots in Education?
Today several industry leaders are utilizing AI chatbots to improve their customer service and to connect with the ever-increasing audiences to stay significant and noticeable. Aside from business, different sectors are additionally sending chatbots, including educational institutes and instructors. Chatbot developers use artificial intelligence and the latest conversational tools to create bots that can speak with students regarding all matters of primary, optional, secondary school and up to university levels. However, AI will not (but may be in next 20 years) swap a student's preferred teacher but can support as a helper to the educator or the means of modern education. AI-driven tools not only recover student interaction and teamwork but also act as a game changer in the advanced Ed-tech world.
A Complete Machine Learning Project Walk-Through in Python
Reading through a data science book or taking a course, it can feel like you have the individual pieces, but don't quite know how to put them together. Taking the next step and solving a complete machine learning problem can be daunting, but preserving and completing a first project will give you the confidence to tackle any data science problem. This series of articles will walk through a complete machine learning solution with a real-world dataset to let you see how all the pieces come together. We'll follow the general machine learning workflow step-by-step: Along the way, we'll see how each step flows into the next and how to specifically implement each part in Python. The complete project is available on GitHub, with the first notebook here. After completing the work, I was offered the job, but then the CTO of the company quit and they weren't able to bring on any new employees. I guess that's how things go on the start-up scene!) The first step before we get coding is to understand the problem we are trying to solve and the available data. In this project, we will work with publicly available building energy data from New York City. The objective is to use the energy data to build a model that can predict the Energy Star Score of a building and interpret the results to find the factors which influence the score. We want to develop a model that is both accurate *-- it can predict the Energy Star Score close to the true value -- and *interpretable -- we can understand the model predictions. Once we know the goal, we can use it to guide our decisions as we dig into the data and build models. Contrary to what most data science courses would have you believe, not every dataset is a perfectly curated group of observations with no missing values or anomalies (looking at you mtcars and iris datasets). Real-world data is messy which means we need to clean and wrangle it into an acceptable format before we can even start the analysis. Data cleaning is an un-glamorous, but necessary part of most actual data science problems.