Education
Robust Educational Dialogue Act Classifiers with Low-Resource and Imbalanced Datasets
Lin, Jionghao, Tan, Wei, Nguyen, Ngoc Dang, Lang, David, Du, Lan, Buntine, Wray, Beare, Richard, Chen, Guanliang, Gasevic, Dragan
Dialogue acts (DAs) can represent conversational actions of tutors or students that take place during tutoring dialogues. Automating the identification of DAs in tutoring dialogues is significant to the design of dialogue-based intelligent tutoring systems. Many prior studies employ machine learning models to classify DAs in tutoring dialogues and invest much effort to optimize the classification accuracy by using limited amounts of training data (i.e., low-resource data scenario). However, beyond the classification accuracy, the robustness of the classifier is also important, which can reflect the capability of the classifier on learning the patterns from different class distributions. We note that many prior studies on classifying educational DAs employ cross entropy (CE) loss to optimize DA classifiers on low-resource data with imbalanced DA distribution. The DA classifiers in these studies tend to prioritize accuracy on the majority class at the expense of the minority class which might not be robust to the data with imbalanced ratios of different DA classes. To optimize the robustness of classifiers on imbalanced class distributions, we propose to optimize the performance of the DA classifier by maximizing the area under the ROC curve (AUC) score (i.e., AUC maximization). Through extensive experiments, our study provides evidence that (i) by maximizing AUC in the training process, the DA classifier achieves significant performance improvement compared to the CE approach under low-resource data, and (ii) AUC maximization approaches can improve the robustness of the DA classifier under different class imbalance ratios.
A CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition
Fan, Ruchao, Chu, Wei, Chang, Peng, Alwan, Abeer
Recently, end-to-end models have been widely used in automatic speech recognition (ASR) systems. Two of the most representative approaches are connectionist temporal classification (CTC) and attention-based encoder-decoder (AED) models. Autoregressive transformers, variants of AED, adopt an autoregressive mechanism for token generation and thus are relatively slow during inference. In this paper, we present a comprehensive study of a CTC Alignment-based Single-Step Non-Autoregressive Transformer (CASS-NAT) for end-to-end ASR. In CASS-NAT, word embeddings in the autoregressive transformer (AT) are substituted with token-level acoustic embeddings (TAE) that are extracted from encoder outputs with the acoustical boundary information offered by the CTC alignment. TAE can be obtained in parallel, resulting in a parallel generation of output tokens. During training, Viterbi-alignment is used for TAE generation, and multiple training strategies are further explored to improve the word error rate (WER) performance. During inference, an error-based alignment sampling method is investigated in depth to reduce the alignment mismatch in the training and testing processes. Experimental results show that the CASS-NAT has a WER that is close to AT on various ASR tasks, while providing a ~24x inference speedup. With and without self-supervised learning, we achieve new state-of-the-art results for non-autoregressive models on several datasets. We also analyze the behavior of the CASS-NAT decoder to explain why it can perform similarly to AT. We find that TAEs have similar functionality to word embeddings for grammatical structures, which might indicate the possibility of learning some semantic information from TAEs without a language model.
Neural Approaches to Entity-Centric Information Extraction
Artificial Intelligence (AI) has huge impact on our daily lives with applications such as voice assistants, facial recognition, chatbots, autonomously driving cars, etc. Natural Language Processing (NLP) is a cross-discipline of AI and Linguistics, dedicated to study the understanding of the text. This is a very challenging area due to unstructured nature of the language, with many ambiguous and corner cases. In this thesis we address a very specific area of NLP that involves the understanding of entities (e.g., names of people, organizations, locations) in text. First, we introduce a radically different, entity-centric view of the information in text. We argue that instead of using individual mentions in text to understand their meaning, we should build applications that would work in terms of entity concepts. Next, we present a more detailed model on how the entity-centric approach can be used for the entity linking task. In our work, we show that this task can be improved by considering performing entity linking at the coreference cluster level rather than each of the mentions individually. In our next work, we further study how information from Knowledge Base entities can be integrated into text. Finally, we analyze the evolution of the entities from the evolving temporal perspective.
Single-round Self-supervised Distributed Learning using Vision Transformer
Park, Sangjoon, Lee, Ik-Jae, Kim, Jun Won, Ye, Jong Chul
Despite the recent success of deep learning in the field of medicine, the issue of data scarcity is exacerbated by concerns about privacy and data ownership. Distributed learning approaches, including federated learning, have been investigated to address these issues. However, they are hindered by the need for cumbersome communication overheads and weaknesses in privacy protection. To tackle these challenges, we propose a self-supervised masked sampling distillation method for the vision transformer. This method can be implemented without continuous communication and can enhance privacy by utilizing a vision transformer-specific encryption technique. We conducted extensive experiments on two different tasks, which demonstrated the effectiveness of our method. We achieved superior performance compared to the existing distributed learning strategy as well as the fine-tuning only baseline. Furthermore, since the self-supervised model created using our proposed method can achieve a general semantic understanding of the image, we demonstrate its potential as a task-agnostic self-supervised foundation model for various downstream tasks, thereby expanding its applicability in the medical domain.
Fundamentals of Machine Learning for Supply Chain
This course will teach you how to leverage the power of Python to understand complicated supply chain datasets. Even if you are not familiar with supply chain fundamentals, the rich data sets that we will use as a canvas will help orient you with several Pythonic tools and best practices for exploratory data analysis (EDA). As such, though all datasets are geared towards supply chain minded professionals, the lessons are easily generalizable to other use cases.
The AI Job That Pays Up to $335K--and You Don't Need a Computer Engineering Background
A new kind of AI job is emerging--and it pays six-figure salaries and doesn't require a degree in computer engineering, or even advanced coding skills. With the rise in generative artificial intelligence, a host of companies are now looking to hire "prompt engineers" who are tasked with training the emerging crop of AI tools to deliver more accurate and relevant responses to the questions real people are likely to pose. Some of these jobs can even pay up to $335,000 a year. Anna Bernstein, a 29-year-old prompt engineer at generative AI firm Copy.ai in New York, is one of the few people already working in this new field. Her role involves writing text-based prompts that she feeds into the back end of AI tools so they can do things such as generate a blog post or sales email with the proper tone and accurate information.
Top 19 Skills You Need to Know in 2023 to Be a Data Scientist - KDnuggets
If you want to be a data scientist in 2023, there are several new skills you should add to your roster, as well as the slew of existing skills you should have already mastered. Part of the problem is job scope creep. Nobody knows what a data scientist is, or what one should do, least of all your future employer. So anything that has data gets stuck in the data science category for you to deal with. You're expected to know how to clean, transform, statistically analyze, visualize, communicate, and predict data.
Cheating with ChatGPT? Students dish on temptations of AI in the classroom
Students at the University of Texas at Austin tell Fox News whether they know or have heard of fellow students using ChatGPT to complete class assignments. AUSTIN, Texas – A majority of college students who spoke with Fox News said they knew or had heard of fellow pupils using ChatGPT for class assignments. "Unfortunately, yes," Riley, an economics major, told Fox News. "I definitely have heard of a couple of people using it for certain things," Piper, a STEM major, said. Recent advances in artificial intelligence technologies, including ChatGPT and Google's Bard, have ignited plagiarism concerns across American schools.
Banning ChatGPT will do more harm than good
If educators actively engage with students about the technology's capabilities and limitations--and work with them to define new academic standards--ChatGPT, and generative AI more broadly, could both democratize and revitalize K–12 education on an unprecedented scale. A bold claim, I know. Few things are as mentally draining as applying to college these days, and as I slaved away at my supplemental essays, the promise of using ChatGPT as a real-time editor was attractive--partly as a potential productivity boost, but mostly as a distraction. I had ChatGPT carefully review my cloying use of semicolons, grade my writing on a 0–10 scale (the results were erratic and maddening)2, and even role-play as an admissions counselor. Its advice was fundamentally incompatible with the creative demands of the modern college essay, and I mostly ignored it.
Sign Language Translation from Instructional Videos
Tarrés, Laia, Gállego, Gerard I., Duarte, Amanda, Torres, Jordi, Giró-i-Nieto, Xavier
The advances in automatic sign language translation (SLT) to spoken languages have been mostly benchmarked with datasets of limited size and restricted domains. Our work advances the state of the art by providing the first baseline results on How2Sign, a large and broad dataset. We train a Transformer over I3D video features, using the reduced BLEU as a reference metric for validation, instead of the widely used BLEU score. We report a result of 8.03 on the BLEU score, and publish the first open-source implementation of its kind to promote further advances.