Education
End-to-End Speech Recognition: A Survey
Prabhavalkar, Rohit, Hori, Takaaki, Sainath, Tara N., Schlüter, Ralf, Watanabe, Shinji
Within components (models, knowledge sources) of an ASR system the classical approach, deep learning has been introduced before coming to a decision. This is in line with Bayes' to acoustic and language modeling. In acoustic modeling, decision rule, which exactly requires a single global decision deep learning replaced Gaussian mixture distributions (hybrid integrating all available knowledge sources. HMM [3], [4]) or augmented the acoustic feature set c) Joint Training: In terms of model training, E2E suggests (nonlinear disciminant/tandem approach [5], [6]). In language estimating all parameters of all components of a model modeling, deep learning replaced count-based approaches [7], jointly using a single objective function that is consistent with [8], [9]. However, when introducing deep learning, the classical the task at hand, which in case of ASR means minimizing the ASR architecture was not yet touched. Classical stateof-the-art expected word error rate. ASR systems today are composed of many separate d) Training Data: Joint training of an integrated model components and knowledge sources, especially speech signal implies using a single kind of training data, which in case preprocessing, methods for robustness w.r.t.
CTRLStruct: Dialogue Structure Learning for Open-Domain Response Generation
Yin, Congchi, Li, Piji, Ren, Zhaochun
Dialogue structure discovery is essential in dialogue generation. Well-structured topic flow can leverage background information and predict future topics to help generate controllable and explainable responses. However, most previous work focused on dialogue structure learning in task-oriented dialogue other than open-domain dialogue which is more complicated and challenging. In this paper, we present a new framework CTRLStruct for dialogue structure learning to effectively explore topic-level dialogue clusters as well as their transitions with unlabelled information. Precisely, dialogue utterances encoded by bi-directional Transformer are further trained through a special designed contrastive learning task to improve representation. Then we perform clustering to utterance-level representations and form topic-level clusters that can be considered as vertices in dialogue structure graph. The edges in the graph indicating transition probability between vertices are calculated by mimicking expert behavior in datasets. Finally, dialogue structure graph is integrated into dialogue model to perform controlled response generation. Experiments on two popular open-domain dialogue datasets show our model can generate more coherent responses compared to some excellent dialogue models, as well as outperform some typical sentence embedding methods in dialogue utterance representation. Code is available in GitHub.
TextWorldExpress: Simulating Text Games at One Million Steps Per Second
Jansen, Peter A., Côté, Marc-Alexandre
Text-based games offer a challenging test bed to evaluate virtual agents at language understanding, multi-step problem-solving, and common-sense reasoning. However, speed is a major limitation of current text-based games, capping at 300 steps per second, mainly due to the use of legacy tooling. In this work we present TextWorldExpress, a high-performance simulator that includes implementations of three common text game benchmarks that increases simulation throughput by approximately three orders of magnitude, reaching over one million steps per second on common desktop hardware. This significantly reduces experiment runtime, enabling billion-step-scale experiments in about one day.
Gaussian Universality of Perceptrons with Random Labels
Gerace, Federica, Krzakala, Florent, Loureiro, Bruno, Stephan, Ludovic, Zdeborová, Lenka
While classical in many theoretical settings - and in particular in statistical physics-inspired works - the assumption of Gaussian i.i.d. input data is often perceived as a strong limitation in the context of statistics and machine learning. In this study, we redeem this line of work in the case of generalized linear classification, a.k.a. the perceptron model, with random labels. We argue that there is a large universality class of high-dimensional input data for which we obtain the same minimum training loss as for Gaussian data with corresponding data covariance. In the limit of vanishing regularization, we further demonstrate that the training loss is independent of the data covariance. On the theoretical side, we prove this universality for an arbitrary mixture of homogeneous Gaussian clouds. Empirically, we show that the universality holds also for a broad range of real datasets.
Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Lu, Pan, Qiu, Liang, Chang, Kai-Wei, Wu, Ying Nian, Zhu, Song-Chun, Rajpurohit, Tanmay, Clark, Peter, Kalyan, Ashwin
Mathematical reasoning, a core ability of human intelligence, presents unique challenges for machines in abstract thinking and logical reasoning. Recent large pre-trained language models such as GPT-3 have achieved remarkable progress on mathematical reasoning tasks written in text form, such as math word problems (MWP). However, it is unknown if the models can handle more complex problems that involve math reasoning over heterogeneous information, such as tabular data. To fill the gap, we present Tabular Math Word Problems (TabMWP), a new dataset containing 38,431 open-domain grade-level problems that require mathematical reasoning on both textual and tabular data. Each question in TabMWP is aligned with a tabular context, which is presented as an image, semi-structured text, and a structured table. There are two types of questions: free-text and multi-choice, and each problem is annotated with gold solutions to reveal the multi-step reasoning process. We evaluate different pre-trained models on TabMWP, including the GPT-3 model in a few-shot setting. As earlier studies suggest, since few-shot GPT-3 relies on the selection of in-context examples, its performance is unstable and can degrade to near chance. The unstable issue is more severe when handling complex problems like TabMWP. To mitigate this, we further propose a novel approach, PromptPG, which utilizes policy gradient to learn to select in-context examples from a small amount of training data and then constructs the corresponding prompt for the test example. Experimental results show that our method outperforms the best baseline by 5.31% on the accuracy metric and reduces the prediction variance significantly compared to random selection, which verifies its effectiveness in selecting in-context examples.
LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval
Zhang, Kai, Tao, Chongyang, Shen, Tao, Xu, Can, Geng, Xiubo, Jiao, Binxing, Jiang, Daxin
Retrieval models based on dense representations in semantic space have become an indispensable branch for first-stage retrieval. These retrievers benefit from surging advances in representation learning towards compressive global sequence-level embeddings. However, they are prone to overlook local salient phrases and entity mentions in texts, which usually play pivot roles in first-stage retrieval. To mitigate this weakness, we propose to make a dense retriever align a well-performing lexicon-aware representation model. The alignment is achieved by weakened knowledge distillations to enlighten the retriever via two aspects -- 1) a lexicon-augmented contrastive objective to challenge the dense encoder and 2) a pair-wise rank-consistent regularization to make dense model's behavior incline to the other. We evaluate our model on three public benchmarks, which shows that with a comparable lexicon-aware retriever as the teacher, our proposed dense one can bring consistent and significant improvements, and even outdo its teacher. In addition, we found our improvement on the dense retriever is complementary to the standard ranker distillation, which can further lift state-of-the-art performance.
ChatGPT wouldn't exist without Canadian AI pioneers. Why one fears for the future
When ChatGPT was released late last year, people around the world suddenly awoke to the major advancements going on in the world of artificial intelligence (AI). For many, what once seemed like a science fiction fantasy was now reality. In truth, the technology behind the groundbreaking chatbot had been brewing behind the scenes in research labs and major tech companies for years. But refined and released in its most accessible form yet, ChatGPT stands to herald in a transformational age of AI adoption. ChatGPT, and other generative AIs like DALL-E, which can create original text and images from a simple prompt, won't just transform education. It will reshape the way people conduct business, create art and do research. Commentators have likened what's coming to the next Industrial Revolution: one in which the role of humans may radically change. While ChatGPT and DALL-E are both products of OpenAI, an American research company, other Silicon Valley giants have been moving quickly to show they're capable of similar technology. With names like OpenAI, Microsoft, Google, Meta and even Baidu capturing international headlines for their generative AI offerings, it's easy to forget that the foundational principles upon which these technologies rest were developed in large part by Canadian scientists.
As battle persists over AI, here's what teachers, students have to say about ChatGPT use.
Despite concerns about whether students are using ChatGPT to cheat on exams or as a shortcut to doing their coursework, a new national survey shows students and teachers have quickly incorporated the new technology into their every day lives. Laila Ayala, a student at Comp Sci High in New York City, for instance, has used ChatGPT to research prompts for her debate team on the effect of AI on students, student mental health and whether the SAT and ACT should be abolished. In Kentucky, high school junior Zachary Clifton said he's used ChatGPT to create study guides for some of the college courses he takes at a nearby community college. Even as some school districts ban the artificial intelligence platform – which can quickly answers questions about nearly any subject it's asked – and some college professors find themselves becoming hypervigilant about whether students are using it to cheat, the new survey commissioned by the Walton Family Foundation and conducted by Impact Research found 22% of students use the chatbot to help them with coursework or in extracurricular activities "on a weekly basis or more." And more than half of teachers surveyed reported using ChatGPT at least once since its release, with 40% of teachers using it "at least once a week."
Banned – David Hopkins / Education & Leadership
So the number of universities banning the use of ChatGPT or other AI-driven tools in assessment is increasing. News in the UK has just picked up on Russell Group universities, including Oxford and Cambridge, banning the tools. This follows news of New York schools doing the same, and Australia too. When something gets banned, it almost makes it more desirable – I remember when music was labelled with the'parental advisory' sticker in the late 80s, instantly making it more desirable and an album I and my friends just had to have. I'm sure I understood the meaning behind the sticker, but it had the reverse effect – it highlighted music that I may not have noticed before, and made it a target for me and others.
The Kaggle Blueprints: Unlocking Winning Approaches to Data Science Competitions
If you ask any successful Kaggler what tips they have to improve your data science skill set, they all have the same answer. They will tell you to study the top solutions of completed Kaggle competitions. Kaggle is a platform for data science competitions for various types of problems. Competitors compete by building Machine Learning models and submitting their predictions. The competitor with the most accurate predictions takes home a prize.