Goto

Collaborating Authors

 Education


EPIK: Eliminating multi-model Pipelines with Knowledge-distillation

arXiv.org Artificial Intelligence

Real-world tasks are largely composed of multiple models, each performing a sub-task in a larger chain of tasks, i.e., using the output from a model as input for another model in a multi-model pipeline. A model like MATRa performs the task of Crosslingual Transliteration in two stages, using English as an intermediate transliteration target when transliterating between two indic languages. We propose a novel distillation technique, EPIK, that condenses two-stage pipelines for hierarchical tasks into a single end-to-end model without compromising performance. This method can create end-to-end models for tasks without needing a dedicated end-to-end dataset, solving the data scarcity problem. The EPIK model has been distilled from the MATra model using this technique of knowledge distillation. The MATra model can perform crosslingual transliteration between 5 languages - English, Hindi, Tamil, Kannada and Bengali. The EPIK model executes the task of transliteration without any intermediate English output while retaining the performance and accuracy of the MATra model. The EPIK model can perform transliteration with an average CER score of 0.015 and average phonetic accuracy of 92.1%. In addition, the average time for execution has reduced by 54.3% as compared to the teacher model and has a similarity score of 97.5% with the teacher encoder. In a few cases, the EPIK model (student model) can outperform the MATra model (teacher model) even though it has been distilled from the MATra model.


Class-aware Information for Logit-based Knowledge Distillation

arXiv.org Artificial Intelligence

Knowledge distillation aims to transfer knowledge to the student model by utilizing the predictions/features of the teacher model, and feature-based distillation has recently shown its superiority over logit-based distillation. However, due to the cumbersome computation and storage of extra feature transformation, the training overhead of feature-based methods is much higher than that of logit-based distillation. In this work, we revisit the logit-based knowledge distillation, and observe that the existing logit-based distillation methods treat the prediction logits only in the instance level, while many other useful semantic information is overlooked. To address this issue, we propose a Class-aware Logit Knowledge Distillation (CLKD) method, that extents the logit distillation in both instance-level and class-level. CLKD enables the student model mimic higher semantic information from the teacher model, hence improving the distillation performance. We further introduce a novel loss called Class Correlation Loss to force the student learn the inherent class-level correlation of the teacher. Empirical comparisons demonstrate the superiority of the proposed method over several prevailing logit-based methods and feature-based methods, in which CLKD achieves compelling results on various visual classification tasks and outperforms the state-of-the-art baselines.


Unbiased Knowledge Distillation for Recommendation

arXiv.org Artificial Intelligence

As a promising solution for model compression, knowledge distillation (KD) has been applied in recommender systems (RS) to reduce inference latency. Traditional solutions first train a full teacher model from the training data, and then transfer its knowledge (\ie \textit{soft labels}) to supervise the learning of a compact student model. However, we find such a standard distillation paradigm would incur serious bias issue -- popular items are more heavily recommended after the distillation. This effect prevents the student model from making accurate and fair recommendations, decreasing the effectiveness of RS. In this work, we identify the origin of the bias in KD -- it roots in the biased soft labels from the teacher, and is further propagated and intensified during the distillation. To rectify this, we propose a new KD method with a stratified distillation strategy. It first partitions items into multiple groups according to their popularity, and then extracts the ranking knowledge within each group to supervise the learning of the student. Our method is simple and teacher-agnostic -- it works on distillation stage without affecting the training of the teacher model. We conduct extensive theoretical and empirical studies to validate the effectiveness of our proposal. We release our code at: https://github.com/chengang95/UnKD.


Deep Neural Networks with PyTorch

#artificialintelligence

The course will teach you how to develop deep learning models using Pytorch. The course will start with Pytorch's tensors and Automatic differentiation package. Then each section will cover different models starting off with fundamentals such as Linear Regression, and logistic/softmax regression. Then Convolutional Neural Networks and Transfer learning will be covered. Finally, several other Deep learning methods will be covered.


End to End Data Science Life Cycle

#artificialintelligence

Information is the oil of the 21st century, and analytics is the combustion engine -- Peter Sondergaard (Senior Vice President and Global Head of Research at Gartner, Inc.) Data science is all about asking interesting questions based on the data you have or often the data you don't have -- Sarah Jarvis (Director of Applied Machine Learning and Data Science at Secondmind) The world we are living in right now is in the era of huge databases. We are living in a digital age where our lifestyle generates more and more data. This data is produced from different sources like Apps, Websites, Smart Devices etc. So, all of this raw data is stored in various Databases. Storing the data doesn't make any sense unless it is used properly for generating insights from the data which helps us to solve various Business problems. With the increasing demand for this field, it is extremely important for us to understand different stages in the life cycle of a Data Science project from End-To-End.


Introduction to Embedded Machine Learning

#artificialintelligence

Machine learning (ML) allows us to teach computers to make predictions and decisions based on data and learn from experiences. In recent years, incredible optimizations have been made to machine learning algorithms, software frameworks, and embedded hardware. Thanks to this, running deep neural networks and other complex machine learning algorithms is possible on low-power devices like microcontrollers. This course will give you a broad overview of how machine learning works, how to train neural networks, and how to deploy those networks to microcontrollers, which is known as embedded machine learning or TinyML. You do not need any prior machine learning knowledge to take this course.


"data science" OR #datascience_2022-11-25_16-06-46.xlsx

#artificialintelligence

The graph represents a network of 3,855 Twitter users whose tweets in the requested range contained ""data science" OR #datascience", or who were replied to or mentioned in those tweets. The network was obtained from the NodeXL Graph Server on Saturday, 26 November 2022 at 00:15 UTC. The requested start date was Friday, 25 November 2022 at 01:01 UTC and the maximum number of days (going backward) was 14. The maximum number of tweets collected was 7,500. The tweets in the network were tweeted over the 3-day, 1-hour, 0-minute period from Monday, 21 November 2022 at 23:59 UTC to Friday, 25 November 2022 at 01:00 UTC.


tensorflow_2022-11-23_03-28-01.xlsx

#artificialintelligence

The graph represents a network of 1,399 Twitter users whose tweets in the requested range contained "tensorflow", or who were replied to or mentioned in those tweets. The network was obtained from the NodeXL Graph Server on Wednesday, 23 November 2022 at 11:35 UTC. The requested start date was Wednesday, 23 November 2022 at 01:01 UTC and the maximum number of days (going backward) was 14. The maximum number of tweets collected was 7,500. The tweets in the network were tweeted over the 2-day, 22-hour, 40-minute period from Saturday, 19 November 2022 at 20:10 UTC to Tuesday, 22 November 2022 at 18:51 UTC.


iot machinelearning_2022-11-23_05-12-01.xlsx

#artificialintelligence

The graph represents a network of 1,680 Twitter users whose tweets in the requested range contained "iot machinelearning", or who were replied to or mentioned in those tweets. The network was obtained from the NodeXL Graph Server on Wednesday, 23 November 2022 at 13:18 UTC. The requested start date was Wednesday, 23 November 2022 at 01:01 UTC and the maximum number of tweets (going backward in time) was 7,500. The tweets in the network were tweeted over the 19-day, 9-hour, 33-minute period from Thursday, 03 November 2022 at 15:26 UTC to Wednesday, 23 November 2022 at 00:59 UTC. Additional tweets that were mentioned in this data set were also collected from prior time periods.


iot bigdata_2022-11-23_04-37-21.xlsx

#artificialintelligence

The graph represents a network of 1,719 Twitter users whose tweets in the requested range contained "iot bigdata", or who were replied to or mentioned in those tweets. The network was obtained from the NodeXL Graph Server on Wednesday, 23 November 2022 at 12:43 UTC. The requested start date was Wednesday, 23 November 2022 at 01:01 UTC and the maximum number of tweets (going backward in time) was 7,500. The tweets in the network were tweeted over the 19-day, 11-hour, 51-minute period from Thursday, 03 November 2022 at 13:08 UTC to Wednesday, 23 November 2022 at 00:59 UTC. Additional tweets that were mentioned in this data set were also collected from prior time periods.