Education
CAVIAR: Categorical-Variable Embeddings for Accurate and Robust Inference
Mukherjee, Anirban, Chang, Hannah Hanwen
Social science research often hinges on the relationship between categorical variables and outcomes. We introduce CAVIAR, a novel method for embedding categorical variables that assume values in a high-dimensional ambient space but are sampled from an underlying manifold. Our theoretical and numerical analyses outline challenges posed by such categorical variables in causal inference. Specifically, dynamically varying and sparse levels can lead to violations of the Donsker conditions and a failure of the estimation functionals to converge to a tight Gaussian process. Traditional approaches, including the exclusion of rare categorical levels and principled variable selection models like LASSO, fall short. CAVIAR embeds the data into a lower-dimensional global coordinate system. The mapping can be derived from both structured and unstructured data, and ensures stable and robust estimates through dimensionality reduction. In a dataset of direct-to-consumer apparel sales, we illustrate how high-dimensional categorical variables, such as zip codes, can be succinctly represented, facilitating inference and analysis.
Best Practices and Lessons Learned on Synthetic Data for Language Models
Liu, Ruibo, Wei, Jerry, Liu, Fangyu, Si, Chenglei, Zhang, Yanzhe, Rao, Jinmeng, Zheng, Steven, Peng, Daiyi, Yang, Diyi, Zhou, Denny, Dai, Andrew M.
The success of AI models relies on the availability of large, diverse, and high-quality datasets, which can be challenging to obtain due to data scarcity, privacy concerns, and high costs. Synthetic data has emerged as a promising solution by generating artificial data that mimics real-world patterns. This paper provides an overview of synthetic data research, discussing its applications, challenges, and future directions. We present empirical evidence from prior art to demonstrate its effectiveness and highlight the importance of ensuring its factuality, fidelity, and unbiasedness. We emphasize the need for responsible use of synthetic data to build more powerful, inclusive, and trustworthy language models.
KTbench: A Novel Data Leakage-Free Framework for Knowledge Tracing
Badran, Yahya, Preisach, Christine
Knowledge Tracing (KT) is concerned with predicting students' future performance on learning items in intelligent tutoring systems. Learning items are tagged with skill labels called knowledge concepts (KCs). Many KT models expand the sequence of item-student interactions into KC-student interactions by replacing learning items with their constituting KCs. This often results in a longer sequence length. This approach addresses the issue of sparse item-student interactions and minimises model parameters. However, two problems have been identified with such models. The first problem is the model's ability to learn correlations between KCs belonging to the same item, which can result in the leakage of ground truth labels and hinder performance. This problem can lead to a significant decrease in performance on datasets with a higher number of KCs per item. The second problem is that the available benchmark implementations ignore accounting for changes in sequence length when expanding KCs, leading to different models being tested with varying sequence lengths but still compared against the same benchmark. To address these problems, we introduce a general masking framework that mitigates the first problem and enhances the performance of such KT models while preserving the original model architecture without significant alterations. Additionally, we introduce KTbench, an open-source benchmark library designed to ensure the reproducibility of this work while mitigating the second problem.
An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution
Lo, Tien-Hong, Chao, Fu-An, Wu, Tzu-I, Sung, Yao-Ting, Chen, Berlin
Automated speaking assessment (ASA) typically involves automatic speech recognition (ASR) and hand-crafted feature extraction from the ASR transcript of a learner's speech. Recently, self-supervised learning (SSL) has shown stellar performance compared to traditional methods. However, SSL-based ASA systems are faced with at least three data-related challenges: limited annotated data, uneven distribution of learner proficiency levels and non-uniform score intervals between different CEFR proficiency levels. To address these challenges, we explore the use of two novel modeling strategies: metric-based classification and loss reweighting, leveraging distinct SSL-based embedding features. Extensive experimental results on the ICNALE benchmark dataset suggest that our approach can outperform existing strong baselines by a sizable margin, achieving a significant improvement of more than 10% in CEFR prediction accuracy.
MultiLS-SP/CA: Lexical Complexity Prediction and Lexical Simplification Resources for Catalan and Spanish
Bott, Stefan, Saggion, Horacio, Rojas, Nelson Perรฉz, Salazar, Martin Solis, Ramirez, Saul Calderon
Automatic lexical simplification is a task to substitute lexical items that may be unfamiliar and difficult to understand with easier and more common words. This paper presents MultiLS-SP/CA, a novel dataset for lexical simplification in Spanish and Catalan. This dataset represents the first of its kind in Catalan and a substantial addition to the sparse data on automatic lexical simplification which is available for Spanish. Specifically, MultiLS-SP is the first dataset for Spanish which includes scalar ratings of the understanding difficulty of lexical items. In addition, we describe experiments with this dataset, which can serve as a baseline for future work on the same data.
Scalable Language Model with Generalized Continual Learning
Peng, Bohao, Tian, Zhuotao, Liu, Shu, Yang, Mingchang, Jia, Jiaya
Continual learning has gained increasing importance as it facilitates the acquisition and refinement of scalable knowledge and skills in language models. However, existing methods typically encounter strict limitations and challenges in real-world scenarios, such as reliance on experience replay, optimization constraints, and inference task-ID. In this study, we introduce the Scalable Language Model (SLM) to overcome these limitations within a more challenging and generalized setting, representing a significant advancement toward practical applications for continual learning. Specifically, we propose the Joint Adaptive Re-Parameterization (JARe), integrated with Dynamic Task-related Knowledge Retrieval (DTKR), to enable adaptive adjustment of language models based on specific downstream tasks. This approach leverages the task distribution within the vector space, aiming to achieve a smooth and effortless continual learning process. Our method demonstrates state-of-the-art performance on diverse backbones and benchmarks, achieving effective continual learning in both full-set and few-shot scenarios with minimal forgetting. Moreover, while prior research primarily focused on a single task type such as classification, our study goes beyond, with the large language model, i.e., LLaMA-2, to explore the effects across diverse domains and task types, such that a single language model can be decently scaled to broader applications.
Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios
Zhang, Yuan, Tao, Xiaomei, Ai, Hanxu, Chen, Tao, Gan, Yanling
In the Massive Open Online Courses (MOOC) learning scenario, the semantic information of instructional videos has a crucial impact on learners' emotional state. Learners mainly acquire knowledge by watching instructional videos, and the semantic information in the videos directly affects learners' emotional states. However, few studies have paid attention to the potential influence of the semantic information of instructional videos on learners' emotional states. To deeply explore the impact of video semantic information on learners' emotions, this paper innovatively proposes a multimodal emotion recognition method by fusing video semantic information and physiological signals. We generate video descriptions through a pre-trained large language model (LLM) to obtain high-level semantic information about instructional videos. Using the cross-attention mechanism for modal interaction, the semantic information is fused with the eye movement and PhotoPlethysmoGraphy (PPG) signals to obtain the features containing the critical information of the three modes. The accurate recognition of learners' emotional states is realized through the emotion classifier. The experimental results show that our method has significantly improved emotion recognition performance, providing a new perspective and efficient method for emotion recognition research in MOOC learning scenarios. The method proposed in this paper not only contributes to a deeper understanding of the impact of instructional videos on learners' emotional states but also provides a beneficial reference for future research on emotion recognition in MOOC learning scenarios.
Why China's regulators are softening on its tech sector
So I was inspired after talking to Angela Huyue Zhang, a law professor in Hong Kong who's coming to teach at the University of Southern California this fall, about her new book on interpreting the logic and patterns behind China's tech regulations. We talked about how the Chinese government almost always swings back and forth between regulating tech too much and not enough, how local governments have gone to great lengths to protect local tech companies, and why AI companies in China are receiving more government goodwill than other sectors today. To learn more about Zhang's fascinating interpretation of the tech regulations in China, read my story published today. In this newsletter, I want to show you a particularly interesting part of the conversation we had, where Zhang expanded on how market overreactions to Chinese tech policies have become an integral part of the tech regulator's toolbox today. The capital markets, perpetually betting on whether tech companies are going to fare better or worse, are always looking for policy signals on whether China is going to start a new crackdown on certain technologies. As a result, they often overreact to every move by the Chinese government.
Microsoft to invest 2.9 billion to boost AI, cloud in Japan
Microsoft will invest 2.9 billion over the next two years to boost its hyperscale cloud computing and artificial intelligence infrastructure in Japan, marking its biggest investment in the country. The announcement was made on Tuesday in Washington after Microsoft President Brad Smith met Prime Minister Fumio Kishida, who is in the United States for the first official visit by a Japanese leader in nine years. The Nikkei newspaper had reported the new investment earlier. Microsoft will also expand its digital training programs to provide AI skills to more than 3 million people over the next three years, the company said in a statement. It plans to open a lab in Japan focused on AI and robotics, while deepening its cybersecurity collaboration with the Japanese government.
L.A. school district probes inappropriate images shared at Fairfax High. More AI abuse?
Los Angeles school officials are investigating allegations that inappropriate photos were "created and disseminated within the Fairfax High School community," in what appears to be the latest alleged misuse of technology by students, a district statement said. Last week, Laguna Beach High School administrators announced that they had launched an investigation after a student allegedly created and circulated "inappropriate images" of classmates through the use of artificial intelligence. In January, five Beverly Hills eighth-graders were expelled for their involvement in the creation and sharing of fake nude pictures of classmates. The students superimposed pictures of classmates' faces onto nude bodies generated by artificial intelligence. In total, 16 eighth-grade students were targeted by the pictures, which were shared through messaging apps, according to the district.