Goto

Collaborating Authors

 Education


Meta Learning as Bayes Risk Minimization

arXiv.org Machine Learning

We show that, when we cast meta-learning problem as BRM, the optimal solution Meta-Learning is a family of methods that use is given by the predictive distribution computed from a set of interrelated tasks to learn a model that the posterior distribution of the latent variable conditioned can quickly learn a new query task from a possibly against the contextual dataset. This result justifies the use of small contextual dataset. In this study, we the predictive distribution in many previous studies of meta use a probabilistic framework to formalize what learning, such as (Edwards & Storkey, 2017; Gordon et al., it means for two tasks to be related and reframe 2018; Garnelo et al., 2018). However, the optimality of the the meta-learning problem into the problem of predictive distribution cannot be guaranteed if one uses an Bayesian risk minimization (BRM). In our formulation, approximation of the posterior distribution that violates the the BRM optimal solution is given by the way the posterior distribution changes with the contextual predictive distribution computed from the posterior dataset, and this is unfortunately the case for most of the distribution of the task-specific latent variable aforementioned works. For example, the variance of the conditioned on the contextual dataset, and this posterior in these works do not converge to 0 as we take justifies the philosophy of Neural Process.


Offline and Online Satisfaction Prediction in Open-Domain Conversational Systems

arXiv.org Artificial Intelligence

Predicting user satisfaction in conversational systems has become critical, as spoken conversational assistants operate in increasingly complex domains. Online satisfaction prediction (i.e., predicting satisfaction of the user with the system after each turn) could be used as a new proxy for implicit user feedback, and offers promising opportunities to create more responsive and effective conversational agents, which adapt to the user's engagement with the agent. To accomplish this goal, we propose a conversational satisfaction prediction model specifically designed for open-domain spoken conversational agents, called ConvSAT. To operate robustly across domains, ConvSAT aggregates multiple representations of the conversation, namely the conversation history, utterance and response content, and system- and user-oriented behavioral signals. We first calibrate ConvSAT performance against state of the art methods on a standard dataset (Dialogue Breakdown Detection Challenge) in an online regime, and then evaluate ConvSAT on a large dataset of conversations with real users, collected as part of the Alexa Prize competition. Our experimental results show that ConvSAT significantly improves satisfaction prediction for both offline and online setting on both datasets, compared to the previously reported state-of-the-art approaches. The insights from our study can enable more intelligent conversational systems, which could adapt in real-time to the inferred user satisfaction and engagement.


IIT-Ropar and TSW Launch a PG Programme in Artificial Intelligence

#artificialintelligence

IIT-Ropar, one of the eight new IITs established by the Ministry of Human Resource Development (MHRD), Government of India, and TSW, the executive education division of Times Professional Learning (a part of The Times of India Group), have launched a Post Graduate Certificate Programme in Artificial Intelligence & Deep Learning. The programme will be coordinated by The Indo-Taiwan Joint Research Centre (ITJRC) on Artificial Intelligence (AI) and Machine Learning (ML), at IIT-Ropar. Supported by the Ministry of Science and Technology, Taiwan, ITJRC is a bilateral centre for collaborative research in disruptive technologies like AI and ML. The programme, with its focus on Artificial Intelligence and Deep Learning, has an eligibility criterion of a minimum of 2 years of work experience in the IT industry. Though an engineering degree is a desirable prerequisite for this programme, one does not need a coding or mathematics background to be eligible.


Mind Blowing Tech in Learning: AI, VR, and AR featuring Prof. Donald Clark @DonaldClark

#artificialintelligence

Hoy traemos a este espacio esta conferencia titulada "Mind Blowing Tech in Learning: AI, VR, and AR featuring" del Prof. Donald Clark, del Center for Online Innovation in Learning y que nos presentan asรญ: Artificial intelligence (AI) is now the most potent force in IT and will shape learning technology, allowing us to escape from the 30 year paradigm of flat, linear e-learning. During this COIL Fischer Speaker Series presentation, Professor Donald Clark debunks some myths about AI and provide real examples of AI used now in content creation, feedback, assessment and spaced practice. In addition he will talk about virtual reality (VR) & augmented reality (AR) as reviving'learning by doing' and their power to democratize experience. Donald Clark is an EdTech entrepreneur and was CEO and one of the original founders of Epic Group plc, which established itself as the leading company in the UK online learning market, floated on the Stock Market in 1996 and sold in 2005, now CEO of Wildfire Ltd. he also invests in, and advises, EdTech companies. Describing himself as'free from the tyranny of employment', he is a board member of Cogbooks, LearningPool, WildFire and Deputy Chair of Brighton Dome & Arts Festival as well as a Visiting Professor at The University of Derby and Fellow of the Royal Society of Arts (FRSA).


3 Levels of Data Science

#artificialintelligence

This article will discuss what I consider to be the three levels of data science competency, namely: level 1 (basic level); level 2 (intermediate level); and level 3 (advanced level). Competency increases from level 1 to 3. We shall use Python as the default language, even though other platforms such as R, SAS, and Matlab could be used as programming languages for data science. The views provided here are my views and are based on my own journey to data science. At level one, a data science aspirant should be able to work with datasets generally presented in comma-separated values (CSV) file format. They should have competency in data basics; data visualization; and linear regression.


Ravi Shankar Prasad launches India's National Artificial Intelligence Portal

#artificialintelligence

New Delhi: The Union Minister for Electronics and IT, Law and Justice and Communications Ravi Shankar Prasad on Saturday (May 30, 2020) launched India s national Artificial Intelligence Portal called www.ai.gov.in. The portal will work as a one stop digital platform for AI related developments in India, sharing of resources such as articles, startups, investment funds in AI, resources, companies and educational institutions related to AI in India. The portal will also share documents, case studies, research reports etc. It has section about learning and new job roles related to AI. This portal has been jointly developed by the Ministry of Electronics and IT and IT Industry.


Govt launches national artificial intelligence mission for industry and schools

#artificialintelligence

The government has launched a new national artificial intelligence (AI) portal which will serve as a knowledge hub for all those who are engaged in this domain. The portal โ€“ www.ai.gov.in was launched by Union Minister for Electronics and IT, Law and Justice and Communications Ravi Shankar Prasad and it was also on the occasion of the first anniversary of the second tenure of the Narendra Modi government. This portal has been jointly developed by the Ministry of Electronics and IT and the IT Industry. The National e-Governance Division of Ministry of Electronics and IT and NASSCOM from the IT industry will jointly run this portal. According to the government, the portal will work as a one-stop digital platform for AI-related developments in India, sharing of resources such as articles, startups, investment funds in AI, resources, companies and educational institutions related to AI in India.


Saber Pro success prediction model using decision tree based learning

arXiv.org Artificial Intelligence

The primary objective of this report is to determine what influences the success rates of students who have studied in Colombia, analyzing the Saber 11, the test done at the last school year, some socioeconomic aspects and comparing the Saber Pro results with the national average. The problem this faces is to find what influences success, but it also provides an insight in the countries education dynamics and predicts one's opportunities to be prosperous. The opposite situation to the one presented in this paper could be the desertion levels, in the sense that by detecting what makes someone outstanding, these factors can say what makes one unsuccessful. The solution proposed to solve this problem was to implement a CART decision tree algorithm that helps to predict the probability that a student has of scoring higher than the mean value, based on different socioeconomic and academic factors, such as the profession of the parents of the subject parents and the results obtained on Saber 11. It was discovered that one of the most influential factors is the score in the Saber 11, on the topic of Social Studies, and that the gender of the subject is not as influential as it is usually portrayed as. The algorithm designed provided significant insight into which factors most affect the probability of success of any given person and if further pursued could be used in many given situations such as deciding which subject in school should be given more intensity to and academic curriculum in general.


ADAHESSIAN: An Adaptive Second Order Optimizer for Machine Learning

arXiv.org Machine Learning

We introduce AdaHessian, a second order stochastic optimization algorithm which dynamically incorporates the curvature of the loss function via ADAptive estimates of the Hessian. Second order algorithms are among the most powerful optimization algorithms with superior convergence properties as compared to first order methods such as SGD and ADAM. The main disadvantage of traditional second order methods is their heavier per-iteration computation and poor accuracy as compared to first order methods. To address these, we incorporate several novel approaches in AdaHessian, including: (i) a new variance reduction estimate of the Hessian diagonal with low computational overhead; (ii) a root-mean-square exponential moving average to smooth out variations of the Hessian diagonal across different iterations; and (iii) a block diagonal averaging to reduce the variance of Hessian diagonal elements. We show that AdaHessian achieves new state-of-the-art results by a large margin as compared to other adaptive optimization methods, including variants of ADAM. In particular, we perform extensive tests on CV, NLP, and recommendation system tasks and find that AdaHessian: (i) achieves 1.80\%/1.45\% higher accuracy on ResNets20/32 on Cifar10, and 5.55\% higher accuracy on ImageNet as compared to ADAM; (ii) outperforms ADAMW for transformers by 0.27/0.33 BLEU score on IWSLT14/WMT14 and 1.8/1.0 PPL on PTB/Wikitext-103; and (iii) achieves 0.032\% better score than AdaGrad for DLRM on the Criteo Ad Kaggle dataset. Importantly, we show that the cost per iteration of AdaHessian is comparable to first-order methods, and that it exhibits robustness towards its hyperparameters. The code for AdaHessian is open-sourced and publicly available.


Semi-supervised deep learning for high-dimensional uncertainty quantification

arXiv.org Machine Learning

This paper presents a semisupervised system responses evaluations, easy-to-evaluate surrogate models learning framework for dimension reduction and have been utilized as substitutes for computationally expensive reliability analysis. An autoencoder is first adopted for mapping simulations or experiments. Popular choices for surrogate the high-dimensional space into a low-dimensional latent space, models in the literature include, support vector machines (SVM) which contains a distinguishable failure surface. Then a deep [4-7], Kriging models [8-10], and artificial neural networks [11-feedforward neural network (DFN) is utilized to learn the 14]. Given a set of training data, surrogate models can be mapping relationship and reconstruct the latent space, while the constructed and then MCS can be directly carried out for Gaussian process (GP) modeling technique is used to build the reliability analysis. Research efforts have been devoted to surrogate model of the transformed limit state function. During developing adaptive sampling strategies [15-18], which aim at the training process of the DFN, the discrepancy between the balancing the fidelity of the surrogate model and the costs of actual and reconstructed latent space is minimized through semisupervised function evaluations.