Goto

Collaborating Authors

 Education


Comparison of Soft and Hard Target RNN-T Distillation for Large-scale ASR

arXiv.org Artificial Intelligence

Knowledge distillation is an effective machine learning technique to transfer knowledge from a teacher model to a smaller student model, especially with unlabeled data. In this paper, we focus on knowledge distillation for the RNN-T model, which is widely used in state-of-the-art (SoTA) automatic speech recognition (ASR). Specifically, we compared using soft and hard target distillation to train large-scaleRNN-T models on the LibriSpeech/LibriLight public dataset (60k hours) and our in-house data (600k hours). We found that hard tar-gets are more effective when the teacher and student have different architecture, such as large teacher and small streaming student. On the other hand, soft target distillation works better in self-training scenario like iterative large teacher training. For a large model with0.6B weights, we achieve a new SoTA word error rate (WER) on LibriSpeech (8% relative improvement on dev-other) using Noisy Student Training with soft target distillation. It also allows our production teacher to adapt new data domain continuously.


Investigating Ensemble Methods for Model Robustness Improvement of Text Classifiers

arXiv.org Artificial Intelligence

Large pre-trained language models have shown remarkable performance over the past few years. These models, however, sometimes learn superficial features from the dataset and cannot generalize to the distributions that are dissimilar to the training scenario. There have been several approaches proposed to reduce model's reliance on these bias features which can improve model robustness in the out-of-distribution setting. However, existing methods usually use a fixed low-capacity model to deal with various bias features, which ignore the learnability of those features. In this paper, we analyze a set of existing bias features and demonstrate there is no single model that works best for all the cases. We further show that by choosing an appropriate bias model, we can obtain a better robustness result than baselines with a more sophisticated model design.


Mapping Husserlian phenomenology onto active inference

arXiv.org Artificial Intelligence

Phenomenology is the rigorous descriptive study of conscious experience. Recent attempts to formalize Husserlian phenomenology provide us with a mathematical model of perception as a function of prior knowledge and expectation. In this paper, we re-examine elements of Husserlian phenomenology through the lens of active inference. In doing so, we aim to advance the project of computational phenomenology, as recently outlined by proponents of active inference. We propose that key aspects of Husserl's descriptions of consciousness can be mapped onto aspects of the generative models associated with the active inference approach. We first briefly review active inference. We then discuss Husserl's phenomenology, with a focus on time consciousness. Finally, we present our mapping from Husserlian phenomenology to active inference.


Visual Answer Localization with Cross-modal Mutual Knowledge Transfer

arXiv.org Artificial Intelligence

The goal of visual answering localization (VAL) in the video is to obtain a relevant and concise time clip from a video as the answer to the given natural language question. Early methods are based on the interaction modelling between video and text to predict the visual answer by the visual predictor. Later, using the textual predictor with subtitles for the VAL proves to be more precise. However, these existing methods still have cross-modal knowledge deviations from visual frames or textual subtitles. In this paper, we propose a cross-modal mutual knowledge transfer span localization (MutualSL) method to reduce the knowledge deviation. MutualSL has both visual predictor and textual predictor, where we expect the prediction results of these both to be consistent, so as to promote semantic knowledge understanding between cross-modalities. On this basis, we design a one-way dynamic loss function to dynamically adjust the proportion of knowledge transfer. We have conducted extensive experiments on three public datasets for evaluation. The experimental results show that our method outperforms other competitive state-of-the-art (SOTA) methods, demonstrating its effectiveness.


AI is changing scientists' understanding of language learning – and raising questions about an innate grammar

#artificialintelligence

Unlike the carefully scripted dialogue found in most books and movies, the language of everyday interaction tends to be messy and incomplete, full of false starts, interruptions and people talking over each other. From casual conversations between friends, to bickering between siblings, to formal discussions in a boardroom, authentic conversation is chaotic. It seems miraculous that anyone can learn language at all given the haphazard nature of the linguistic experience. For this reason, many language scientists – including Noam Chomsky, a founder of modern linguistics – believe that language learners require a kind of glue to rein in the unruly nature of everyday language. And that glue is grammar: a system of rules for generating grammatical sentences.


Supervised Machine Learning: Regression

#artificialintelligence

This course introduces you to one of the main types of modelling families of supervised Machine Learning: Regression. You will learn how to train regression models to predict continuous outcomes and how to use error metrics to compare across different models. This course also walks you through best practices, including train and test splits, and regularization techniques. By the end of this course you should be able to: Differentiate uses and applications of classification and regression in the context of supervised machine learning Describe and use linear regression models Use a variety of error metrics to compare and select a linear regression model that best suits your data Articulate why regularization may help prevent overfitting Use regularization regressions: Ridge, LASSO, and Elastic net Who should take this course? This course targets aspiring data scientists interested in acquiring hands-on experience with Supervised Machine Learning Regression techniques in a business setting.


AIhub monthly digest: October 2022 – Nigerian sign language, a simple voting rule, and robotic control algorithms

AIHub

Welcome to our October 2022 monthly digest, where you can catch up with any AIhub stories you may have missed, get the low-down on recent events, and much more. This month, we learn about a Nigerian sign language dataset, hear from researchers working on different robotic control projects, and dig into the latest governmental AI policies. Steven Kolawole created a pioneering dataset for Nigerian sign language, in collaboration with a TV sign language broadcaster and two schools in Nigeria. He used this dataset of over 8000 images to create a model to convert sign language to text or speech. In this interview, Steven told us about the goals of this research, his methodology, and how the work has inspired research in other languages.


Best Udacity Nanodegree for Machine learning You Should Enroll in 2022

#artificialintelligence

I hope you have found the Best Udacity Nanodegree for Machine learning. I would suggest you bookmark this article for future referrals. Now it's time to wrap up. In this article, I tried to cover the Best Udacity Nanodegree Programs for Machine learning. If you have any doubts or questions, feel free to ask me in the comment section. Best Math Courses for Machine Learning- Find the Best One! 9 Best Tensorflow Courses & Certifications Online- Discover the Best One! Machine Learning Engineer Career Path: Step by Step Complete Guide Best Online Courses On Machine Learning You Must Know in 2022 Best Machine Learning Courses for Finance You Must Know Best Resources to Learn Machine Learning Online in 2022


A Question-Answering Bot Powered by Wikipedia, Coupled to GPT-3

#artificialintelligence

If you follow me, you've seen I'm fascinated with GPT-3 both as a tool for productivity and as a tool for information retrieval through natural questions. You've also seen that GPT-3 often provides correct answers to a question, but sometimes it does not and it can even be misleading or confusing because its answer appears confident despite being wrong. In some cases, but not always, when it cannot find a reasonable completion (i.e. it "doesn't know" the answer) it tells you so, or it just doesn't provide any answer. I showed you that factual accuracy can be improved by fine-tuning the model, or more easily, by few-shot learning. But it isn't easy to decide what information to use in these procedures, let alone how to apply it.


25 Recent News Bites About Biden's Pants, Artificial Intelligence, And A Radioactive Elementary School

#artificialintelligence

In the wake of the Supreme Court's recent decision to reverse Roe v Wade, more and more people across the United States are now looking for vasectomies. This surge in demand is due to people wanting to take matters into their own hands when it comes to preventing pregnancy – and one mobile clinic, called the "Nutcracker", even has a program offering 60 free snips over three days! In only slightly more odd news, Rick Caruso, a billionaire real estate developer and candidate for Los Angeles mayor, recently made headlines after declaring that he's not white - he's actually Italian. This claim came during an awkward debate moment. And, in only slightly less slimy matters, it turns out here may be a new way to protect against HIV - personal lubricant made from cow mucus!