Education
Runtime-Safety-Guided Policy Repair
Zhou, Weichao, Gao, Ruihan, Kim, BaekGyu, Kang, Eunsuk, Li, Wenchao
We study the problem of policy repair for learning-based control policies in safety-critical settings. We consider an architecture where a high-performance learning-based control policy (e.g. one trained as a neural network) is paired with a model-based safety controller. The safety controller is endowed with the abilities to predict whether the trained policy will lead the system to an unsafe state, and take over control when necessary. While this architecture can provide added safety assurances, intermittent and frequent switching between the trained policy and the safety controller can result in undesirable behaviors and reduced performance. We propose to reduce or even eliminate control switching by `repairing' the trained policy based on runtime data produced by the safety controller in a way that deviates minimally from the original policy. The key idea behind our approach is the formulation of a trajectory optimization problem that allows the joint reasoning of policy update and safety constraints. Experimental results demonstrate that our approach is effective even when the system model in the safety controller is unknown and only approximated.
Estimating Causal Effects with the Neural Autoregressive Density Estimator
Garrido, Sergio, Borysov, Stanislav S., Rich, Jeppe, Pereira, Francisco C.
Estimation of causal effects is fundamental in situations were the underlying system will be subject to active interventions. Part of building a causal inference engine is defining how variables relate to each other, that is, defining the functional relationship between variables given conditional dependencies. In this paper, we deviate from the common assumption of linear relationships in causal models by making use of neural autoregressive density estimators and use them to estimate causal effects within the Pearl's do-calculus framework. Using synthetic data, we show that the approach can retrieve causal effects from non-linear systems without explicitly modeling the interactions between the variables.
On the Sample Complexity of Reinforcement Learning with Policy Space Generalization
Mou, Wenlong, Wen, Zheng, Chen, Xi
We study the optimal sample complexity in large-scale Reinforcement Learning (RL) problems with policy space generalization, i.e. the agent has a prior knowledge that the optimal policy lies in a known policy space. Existing results show that without a generalization model, the sample complexity of an RL algorithm will inevitably depend on the cardinalities of state space and action space, which are intractably large in many practical problems. To avoid such undesirable dependence on the state and action space sizes, this paper proposes a new notion of eluder dimension for the policy space, which characterizes the intrinsic complexity of policy learning in an arbitrary Markov Decision Process (MDP). Using a simulator oracle, we prove a near-optimal sample complexity upper bound that only depends linearly on the eluder dimension. We further prove a similar regret bound in deterministic systems without the simulator.
COLD: Towards the Next Generation of Pre-Ranking System
Wang, Zhe, Zhao, Liqin, Jiang, Biye, Zhou, Guorui, Zhu, Xiaoqiang, Gai, Kun
Multi-stage cascade architecture exists widely in many industrial systems such as recommender systems and online advertising, which often consists of sequential modules including matching, pre-ranking, ranking, etc. For a long time, it is believed pre-ranking is just a simplified version of the ranking module, considering the larger size of the candidate set to be ranked. Thus, efforts are made mostly on simplifying ranking model to handle the explosion of computing power for online inference. In this paper, we rethink the challenge of the pre-ranking system from an algorithm-system co-design view. Instead of saving computing power with restriction of model architecture which causes loss of model performance, here we design a new pre-ranking system by joint optimization of both the pre-ranking model and the computing power it costs. We name it COLD (Computing power cost-aware Online and Lightweight Deep pre-ranking system). COLD beats SOTA in three folds: (i) an arbitrary deep model with cross features can be applied in COLD under a constraint of controllable computing power cost. (ii) computing power cost is explicitly reduced by applying optimization tricks for inference acceleration. This further brings space for COLD to apply more complex deep models to reach better performance. (iii) COLD model works in an online learning and severing manner, bringing it excellent ability to handle the challenge of the data distribution shift. Meanwhile, the fully online pre-ranking system of COLD provides us with a flexible infrastructure that supports efficient new model developing and online A/B testing.Since 2019, COLD has been deployed in almost all products involving the pre-ranking module in the display advertising system in Alibaba, bringing significant improvements.
Gradient-based Learning Methods Extended to Smooth Manifolds Applied to Automated Clustering
Koudounas, Alkis, Fiori, Simone
Grassmann manifold based sparse spectral clustering is a classification technique thatย consists in learning a latent representation of data, formed by a subspace basis, whichย is sparse. In order to learn a latent representation, spectral clustering is formulated inย terms of a loss minimization problem over a smooth manifold known as Grassmannian.ย Such minimization problem cannot be tackled by one of traditional gradient-based learningย algorithms, which are only suitable to perform optimization in absence of constraints amongย parameters. It is, therefore, necessary to develop specific optimization/learning algorithmsย that are able to look for a local minimum of a loss function under smooth constraints inย an efficient way. Such need calls for manifold optimization methods. In this paper, weย extend classical gradient-based learning algorithms on ย ย at parameter spaces (from classicalย gradient descent to adaptive momentum) to curved spaces (smooth manifolds) by meansย of tools from manifold calculus. We compare clustering performances of these methodsย and known methods from the scientific literature. The obtained results confirm that theย proposed learning algorithms prove lighter in computational complexity than existing onesย without detriment in clustering efficacy.
Shifu2: A Network Representation Learning Based Model for Advisor-advisee Relationship Mining
Liu, Jiaying, Xia, Feng, Wang, Lei, Xu, Bo, Kong, Xiangjie, Tong, Hanghang, King, Irwin
The advisor-advisee relationship represents direct knowledge heritage, and such relationship may not be readily available from academic libraries and search engines. This work aims to discover advisor-advisee relationships hidden behind scientific collaboration networks. For this purpose, we propose a novel model based on Network Representation Learning (NRL), namely Shifu2, which takes the collaboration network as input and the identified advisor-advisee relationship as output. In contrast to existing NRL models, Shifu2 considers not only the network structure but also the semantic information of nodes and edges. Shifu2 encodes nodes and edges into low-dimensional vectors respectively, both of which are then utilized to identify advisor-advisee relationships. Experimental results illustrate improved stability and effectiveness of the proposed model over state-of-the-art methods. In addition, we generate a large-scale academic genealogy dataset by taking advantage of Shifu2.
Go Wide, Then Narrow: Efficient Training of Deep Thin Networks
Zhou, Denny, Ye, Mao, Chen, Chen, Meng, Tianjian, Tan, Mingxing, Song, Xiaodan, Le, Quoc, Liu, Qiang, Schuurmans, Dale
For deploying a deep learning model into production, it needs to be both accurate and compact to meet the latency and memory constraints. This usually results in a network that is deep (to ensure performance) and yet thin (to improve computational efficiency). In this paper, we propose an efficient method to train a deep thin network with a theoretic guarantee. Our method is motivated by model compression. It consists of three stages. First, we sufficiently widen the deep thin network and train it until convergence. Then, we use this well-trained deep wide network to warm up (or initialize) the original deep thin network. This is achieved by layerwise imitation, that is, forcing the thin network to mimic the intermediate outputs of the wide network from layer to layer. Finally, we further fine tune this already well-initialized deep thin network. The theoretical guarantee is established by using the neural mean field analysis. It demonstrates the advantage of our layerwise imitation approach over backpropagation. We also conduct large-scale empirical experiments to validate the proposed method. By training with our method, ResNet50 can outperform ResNet101, and BERT Base can be comparable with BERT Large, when ResNet101 and BERT Large are trained under the standard training procedures as in the literature.
Cultivating Digital Leaders -- Campus Technology
A program at Georgia State University is widening the scope of digital literacy to include problem-solving and leadership strategies. Most every college or university recognizes the importance of digital literacy to future graduates, and increasingly, you'll find technology competencies and digital skill development experiences included widely in general education programs. But at Georgia State University, leaders within the Center for Excellence in Teaching and Learning are making certain that GSU's digital literacy offerings will spawn not only technically competent individuals, but also a diverse range of professional leaders who know how to use their technology skills in context for better problem solving. Digital Learners to Leaders is an experiential learning program aimed at developing the next generation of digital problem solvers through industry/higher-education partnerships and exposure to digital technologies and the Internet of Things. DLL began as a co-curricular program, seeing its first cohort of 45 students in Spring 2018.
DBS to train staff in artificial intelligence, machine learning via Amazon tie-up
DBS is collaborating with Amazon Web Services (AWS) to equip at least 3,000 of the bank's employees - including senior leadership - with new artificial intelligence (AI) and machine learning (ML) skills by the end of this year, through gamified learning. This is meant to accelerate the use of AI and ML across its business, the bank said in a press statement on Monday. Both companies jointly launched the DBS x AWS DeepRacer League, which will enable DBS employees to learn the basics of AI and ML through a series of hands-on online tutorials. They will then put their new knowledge to the test by programming their own autonomous model race car. These ML models will be uploaded onto a virtual racing environment where the employees can experiment and iteratively finetune their models as they engage each other in friendly competition.
The Surprising Benefits Of AI-Driven Video Conferencing In Education
Artificial intelligence is having a tremendous influence on the future of education. BuiltIn recently published a list of 12 AI startups that specialize in serving the education sector. AI is going to affect education in a number of ways. One of the impacts is the growing use of video conferencing. This year has brought with it a mountain of challenges for educators and parents across the nation.