Education
My 2021 in Review
I look back at my accomplishments in 2021. I worked on various applied ML projects from climate change, protecting sharks to art & design. Highlights include open-source projects, TensorFlow, computer vision, deep learning, art and design. As always, I have been busy learning and sharing my knowledge with the community! Early 2022 I worked as an ML researcher on a Frontier Development Lab project using ML for climate change (MLCC).
Models Are Rarely Deployed: An Industry-wide Failure in Machine Learning Leadership - KDnuggets
The latest KDnuggets poll reconfirms today's dire industry buzz: Very few machine learning models actually get deployed. In this article, I'll summarize the poll results and argue that this pervasive failure of ML projects comes from a lack of prudent leadership. I'll also argue that MLops is not the fundamental missing ingredient โ instead, an effective ML leadership practice must be the dog that wags the model-integration tail. Considering the growing chatter about ML's failure to launch, there's been relatively little concrete industry research โ especially when it comes to surveys on model deployment in particular rather than ROI in general โ so I proposed this poll to Gregory Piatetsky and Matthew Mayo of KDnuggets. They helped me formulate and wordsmith it.
Dynamic Rectification Knowledge Distillation
Amik, Fahad Rahman, Tasin, Ahnaf Ismat, Ahmed, Silvia, Elahi, M. M. Lutfe, Mohammed, Nabeel
Knowledge Distillation is a technique which aims to utilize dark knowledge to compress and transfer information from a vast, well-trained neural network (teacher model) to a smaller, less capable neural network (student model) with improved inference efficiency. This approach of distilling knowledge has gained popularity as a result of the prohibitively complicated nature of such cumbersome models for deployment on edge computing devices. Generally, the teacher models used to teach smaller student models are cumbersome in nature and expensive to train. To eliminate the necessity for a cumbersome teacher model completely, we propose a simple yet effective knowledge distillation framework that we termed Dynamic Rectification Knowledge Distillation (DR-KD). Our method transforms the student into its own teacher, and if the self-teacher makes wrong predictions while distilling information, the error is rectified prior to the knowledge being distilled. Specifically, the teacher targets are dynamically tweaked by the agency of ground-truth while distilling the knowledge gained from traditional training. Our proposed DR-KD performs remarkably well in the absence of a sophisticated cumbersome teacher model and achieves comparable performance to existing state-of-the-art teacher-free knowledge distillation frameworks when implemented by a low-cost dynamic mannered teacher. Our approach is all-encompassing and can be utilized for any deep neural network training that requires categorization or object recognition. DR-KD enhances the test accuracy on Tiny ImageNet by 2.65% over prominent baseline models, which is significantly better than any other knowledge distillation approach while requiring no additional training costs.
Gap Minimization for Knowledge Sharing and Transfer
Wang, Boyu, Mendez, Jorge, Shui, Changjian, Zhou, Fan, Wu, Di, Gagnรฉ, Christian, Eaton, Eric
Learning from multiple related tasks by knowledge sharing and transfer has become increasingly relevant over the last two decades. In order to successfully transfer information from one task to another, it is critical to understand the similarities and differences between the domains. In this paper, we introduce the notion of \emph{performance gap}, an intuitive and novel measure of the distance between learning tasks. Unlike existing measures which are used as tools to bound the difference of expected risks between tasks (e.g., $\mathcal{H}$-divergence or discrepancy distance), we theoretically show that the performance gap can be viewed as a data- and algorithm-dependent regularizer, which controls the model complexity and leads to finer guarantees. More importantly, it also provides new insights and motivates a novel principle for designing strategies for knowledge sharing and transfer: gap minimization. We instantiate this principle with two algorithms: 1. {gapBoost}, a novel and principled boosting algorithm that explicitly minimizes the performance gap between source and target domains for transfer learning; and 2. {gapMTNN}, a representation learning algorithm that reformulates gap minimization as semantic conditional matching for multitask learning. Our extensive evaluation on both transfer learning and multitask learning benchmark data sets shows that our methods outperform existing baselines.
Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection
Gururangan, Suchin, Card, Dallas, Dreier, Sarah K., Gade, Emily K., Wang, Leroy Z., Wang, Zeyu, Zettlemoyer, Luke, Smith, Noah A.
Language models increasingly rely on massive web dumps for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and newswire often serve as anchors for automatically selecting web text most suitable for language modeling, a process typically referred to as quality filtering. Using a new dataset of U.S. high school newspaper articles -- written by students from across the country -- we investigate whose language is preferred by the quality filter used for GPT-3. We find that newspapers from larger schools, located in wealthier, educated, and urban ZIP codes are more likely to be classified as high quality. We then demonstrate that the filter's measurement of quality is unaligned with other sensible metrics, such as factuality or literary acclaim. We argue that privileging any corpus as high quality entails a language ideology, and more care is needed to construct training corpora for language models, with better transparency and justification for the inclusion or exclusion of various texts.
Visualizing the diversity of representations learned by Bayesian neural networks
Grinwald, Dennis, Bykov, Kirill, Nakajima, Shinichi, Hรถhne, Marina M. -C.
Explainable artificial intelligence (XAI) aims to make learning machines less opaque, and offers researchers and practitioners various tools to reveal the decision-making strategies of neural networks. In this work, we investigate how XAI methods can be used for exploring and visualizing the diversity of feature representations learned by Bayesian neural networks (BNNs). Our goal is to provide a global understanding of BNNs by making their decision-making strategies a) visible and tangible through feature visualizations and b) quantitatively measurable with a distance measure learned by contrastive learning. Our work provides new insights into the posterior distribution in terms of human-understandable feature information with regard to the underlying decision-making strategies. Our main findings are the following: 1) global XAI methods can be applied to explain the diversity of decision-making strategies of BNN instances, 2) Monte Carlo dropout exhibits increased diversity in feature representations compared to the multimodal posterior approximation of MultiSWAG, 3) the diversity of learned feature representations highly correlates with the uncertainty estimates, and 4) the inter-mode diversity of the multimodal posterior decreases as the network width increases, while the intra-mode diversity increases. Our findings are consistent with the recent deep neural networks theory, providing additional intuitions about what the theory implies in terms of humanly understandable concepts.
PyTorch for Deep Learning with Python Bootcamp
Welcome to the best online course for learning about Deep Learning with Python and PyTorch! PyTorch is an open source deep learning platform that provides a seamless path from research prototyping to production deployment. It is rapidly becoming one of the most popular deep learning frameworks for Python. Deep integration into Python allows popular libraries and packages to be used for easily writing neural network layers in Python. A rich ecosystem of tools and libraries extends PyTorch and supports development in computer vision, NLP and more.
Question Generation using Natural Language processing
This course focuses on using state-of-the-art Natural Language processing techniques to solve the problem of question generation in edtech. If we pick up any middle school textbook, at the end of every chapter we see assessment questions like MCQs, True/False questions, Fill-in-the-blanks, Match the following, etc. In this course, we will see how we can take any text content and generate these assessment questions using NLP techniques. This course will be a very practical use case of NLP where we put basic algorithms like word vectors (word2vec, Glove, etc) to recent advancements like BERT, openAI GPT-2, and T5 transformers to real-world use. We will use NLP libraries like Spacy, NLTK, AllenNLP, HuggingFace transformers, etc.
Data Analyst
About Apptegy Since our start in 2015, we've gone from a group of individuals to a community pushing toward the same goal of building a fantastic company with great people, great products, and most importantly, a great culture. To date, we have grown from a handful of school districts in Arkansas to thousands of school districts across the U.S. Apptegy is building products to empower school leaders to run better schools. We have the opportunity to help schools as they go through a radical shift in how they operate and to provide great technology to make that transition. Our Engineering team has grown significantly and so too has the number of schools and users. We look forward to meeting with you and telling you more about this opportunity to be part of a growing company, engineering organization, and to support a fast-scaling set of products.
Supervised vs Unsupervised Machine Learning
Artificial intelligence (AI) is being used to change our lives everyday. When it comes to building AI programs, there are two approaches programmers tend to choose: supervised or unsupervised machine learning. The simple distinction between these is supervised machine learning utilizes labeled data to predict outcomes, while unsupervised machine learning does not. There are, however, some differences between the two techniques, as well as critical areas where one surpasses the other. In this article, we will break down some of these differences with examples of both supervised and unsupervised learning.