Education
Efficient Machine Learning
NEW, 4.1 (12 ratings), Created by Usama Albaghdady, English If you're a machine learning specialist looking to make the transaction into the real-world AI applications. This comprehensive course will be your guide to learning how to scale-up your machine learning model to the optimal state possible, you'll be learning everything you need to move you machine learning model to the next stage. This course is designed for both beginners with some programming experience or experienced developers looking to make the jump to Data Science! You'll learn the machine learning, AI, and data mining techniques real employers are looking for, including:
Collaborative Distillation for Top-N Recommendation
Lee, Jae-woong, Choi, Minjin, Lee, Jongwuk, Shim, Hyunjung
--Knowledge distillation (KD) is a well-known method to reduce inference latency by compressing a cumbersome teacher model to a small student model. Despite the success of KD in the classification task, applying KD to recommender models is challenging due to the sparsity of positive feedback, the ambiguity of missing feedback, and the ranking problem associated with the top-N recommendation. T o address the issues, we propose a new KD model for the collaborative filtering approach, namely collaborative distillation ( CD). Specifically, (1) we reformulate a loss function to deal with the ambiguity of missing feedback. Via experimental results, we demonstrate that the proposed model outperforms the state-of-the-art method by 2.7-33.2% Moreover, the proposed model achieves the performance comparable to the teacher model. Neural recommender models [1]-[9] have achieved better performance than conventional latent factor models either by capturing nonlinear and complex correlation patterns among users/items, or by leveraging the hidden features extracted from auxiliary information such as texts and images. However, the number of model parameters of neural models is greater than that of conventional models by one or more orders of magnitude. This indicates a tradeoff between accuracy and efficiency. As a result, neural recommender models usually suffer from higher latency during the inference phase. Our primary goal is to develop a recommender model that achieves a balance between effectiveness and efficiency. In this paper, we employ knowledge distillation (KD) [10] which is a network compression technique by transferring the distilled knowledge of a large model (a.k.a., a teacher model) to a small model (a.k.a., a student model). As the student model can utilize the knowledge transferred from the teacher model, it naturally exhibits the properties of computational efficiency and low memory usage. Therefore, it is capable of achieving a balance between effectiveness and efficiency. Specifically, the training procedure for KD consists of two steps. In the offline training phase, the teacher model is supervised by a training dataset with labels.
Model-Based Reinforcement Learning with Adversarial Training for Online Recommendation
Bai, Xueying, Guan, Jian, Wang, Hongning
Reinforcement learning is effective in optimizing policies for recommender systems. Current solutions mostly focus on model-free approaches, which require frequent interactions with a real environment, and thus are expensive in model learning. Offline evaluation methods, such as importance sampling, can alleviate such limitations, but usually request a large amount of logged data and do not work well when the action space is large. In this work, we propose a model-based reinforcement learning solution which models the user-agent interaction for offline policy learning via a generative adversarial network. To reduce bias in the learnt policy, we use the discriminator to evaluate the quality of generated sequences and rescale the generated rewards. Our theoretical analysis and empirical evaluations demonstrate the effectiveness of our solution in identifying patterns from given offline data and learning policies based on the offline and generated data.
Learning from a Teacher using Unlabeled Data
Menghani, Gaurav, Ravi, Sujith
Knowledge distillation is a widely used technique for model compression. We posit that the teacher model used in a distillation setup, captures relationships between classes, that extend beyond the original dataset. We empirically show that a teacher model can transfer this knowledge to a student model even on an {\it out-of-distribution} dataset. Using this approach, we show promising results on MNIST, CIFAR-10, and Caltech-256 datasets using unlabeled image data from different sources. Our results are encouraging and help shed further light from the perspective of understanding knowledge distillation and utilizing unlabeled data to improve model quality.
Robustness to Capitalization Errors in Named Entity Recognition
Bodapati, Sravan, Yun, Hyokun, Al-Onaizan, Yaser
Robustness to capitalization errors is a highly desirable characteristic of named entity recognizers, yet we find standard models for the task are surprisingly brittle to such noise. Existing methods to improve robustness to the noise completely discard given orthographic information, mwhich significantly degrades their performance on well-formed text. We propose a simple alternative approach based on data augmentation, which allows the model to \emph{learn} to utilize or ignore orthographic information depending on its usefulness in the context. It achieves competitive robustness to capitalization errors while making negligible compromise to its performance on well-formed text and significantly improving generalization power on noisy user-generated text. Our experiments clearly and consistently validate our claim across different types of machine learning models, languages, and dataset sizes.
EDUQA: Educational Domain Question Answering System using Conceptual Network Mapping
Agarwal, Abhishek, Sachdeva, Nikhil, Yadav, Raj Kamal, Udandarao, Vishaal, Mittal, Vrinda, Gupta, Anubha, Mathur, Abhinav
Most of the existing question answering models can be largely compiled into two categories: i) open domain question answering models that answer generic questions and use large-scale knowledge base along with the targeted web-corpus retrieval and ii) closed domain question answering models that address focused questioning area and use complex deep learning models. Both the above models derive answers through textual comprehension methods. Due to their inability to capture the pedagogical meaning of textual content, these models are not appropriately suited to the educational field for pedagogy. In this paper, we propose an on-the-fly conceptual network model that incorporates educational semantics. The proposed model preserves correlations between conceptual entities by applying intelligent indexing algorithms on the concept network so as to improve answer generation. This model can be utilized for building interactive conversational agents for aiding classroom learning.
Machine Intelligence at the Edge with Learning Centric Power Allocation
Wang, Shuai, Wu, Yik-Chung, Xia, Minghua, Wang, Rui, Poor, H. Vincent
While machine-type communication (MTC) devices generate considerable amounts of data, they often cannot process the data due to limited energy and computation power. To empower MTC with intelligence, edge machine learning has been proposed. However, power allocation in this paradigm requires maximizing the learning performance instead of the communication throughput, for which the celebrated water-filling and max-min fairness algorithms become inefficient. To this end, this paper proposes learning centric power allocation (LCPA), which provides a new perspective to radio resource allocation in learning driven scenarios. By employing an empirical classification error model that is supported by learning theory, the LCPA is formulated as a nonconvex nonsmooth optimization problem, and is solved by majorization minimization (MM) framework. To get deeper insights into LCPA, asymptotic analysis shows that the transmit powers are inversely proportional to the channel gain, and scale exponentially with the learning parameters. This is in contrast to traditional power allocations where quality of wireless channels is the only consideration. Last but not least, to enable LCPA in large-scale settings, two optimization algorithms, termed mirror-prox LCPA and accelerated LCPA, are further proposed. Extensive numerical results demonstrate that the proposed LCPA algorithms outperform traditional power allocation algorithms, and the large-scale algorithms reduce the computation time by orders of magnitude compared with MM-based LCPA but still achieve competing learning performance.
Scientists believe programming AI for self-preservation could be the key to giving robots feelings
A new paper from researchers at the University of Southern California's Brain and Creativity Institute considers a novel path toward creating robots with'feelings.' The key, according to researchers Kinson Man and Antonio Damasio, is homestasis, a self-preservation principle by which living creatures seek to maintain internal biological equilibrium by avoiding certain environments or kinds of stimuli. Were robots to be programmed with a homeostatic sense of self-preservation, would that put them on a path toward developing true feelings? According to a Science News report on the paper, Man and Damasio consider the most promising lead for feeling robots to come through the combination of soft robotics and deep learning, which when combined might approximate a homeostatic reaction to negative environmental stimuli. Man and Domasio point to a 1954 experiment by W. Ross Ashby that demonstrated how homeostatic sensing might be translated into robotics.
Master machine learning and AI with this masterclass bundle
The machines are taking over the world! Well, maybe not quite yet, but AI and machine learning are already running much more of the world than you might realize. You could launch a career pioneering these tech innovations with the Machine Learning and Artificial Intelligence Certification Bundle. Today, you can sign up for only $29. Machine learning is the future, and careers are coming with it.
Integrating AI within your Enterprise
Think back to school and science class: You probably conducted an experiment where you placed an alarm clock (set to go off in 5 minutes) under a glass jar and the teacher pumped all the air out of the jar. When the alarm went off, you couldn't hear it, right? Without air, sound does not travel – nature abhors a vacuum. This is also true of artificial intelligence - it cannot survive in a vacuum and needs a rich ecosystem of data where it can thrive. This can only be achieved by integrating "trustworthy AI" systems with the rest of an organization's IT landscape.