Education
Multimodal Methods for Analyzing Learning and Training Environments: A Systematic Literature Review
Cohn, Clayton, Davalos, Eduardo, Vatral, Caleb, Fonteles, Joyce Horn, Wang, Hanchen David, Ma, Meiyi, Biswas, Gautam
Recent technological advancements have enhanced our ability to collect and analyze rich multimodal data (e.g., speech, video, and eye gaze) to better inform learning and training experiences. While previous reviews have focused on parts of the multimodal pipeline (e.g., conceptual models and data fusion), a comprehensive literature review on the methods informing multimodal learning and training environments has not been conducted. This literature review provides an in-depth analysis of research methods in these environments, proposing a taxonomy and framework that encapsulates recent methodological advances in this field and characterizes the multimodal domain in terms of five modality groups: Natural Language, Video, Sensors, Human-Centered, and Environment Logs. We introduce a novel data fusion category -- mid fusion -- and a graph-based technique for refining literature reviews, termed citation graph pruning. Our analysis reveals that leveraging multiple modalities offers a more holistic understanding of the behaviors and outcomes of learners and trainees. Even when multimodality does not enhance predictive accuracy, it often uncovers patterns that contextualize and elucidate unimodal data, revealing subtleties that a single modality may miss. However, there remains a need for further research to bridge the divide between multimodal learning and training studies and foundational AI research.
Diffusion-Based Visual Art Creation: A Survey and New Perspectives
Wang, Bingyuan, Chen, Qifeng, Wang, Zeyu
The integration of generative AI in visual art has revolutionized not only how visual content is created but also how AI interacts with and reflects the underlying domain knowledge. This survey explores the emerging realm of diffusion-based visual art creation, examining its development from both artistic and technical perspectives. We structure the survey into three phases, data feature and framework identification, detailed analyses using a structured coding process, and open-ended prospective outlooks. Our findings reveal how artistic requirements are transformed into technical challenges and highlight the design and application of diffusion-based methods within visual art creation. We also provide insights into future directions from technical and synergistic perspectives, suggesting that the confluence of generative AI and art has shifted the creative paradigm and opened up new possibilities. By summarizing the development and trends of this emerging interdisciplinary area, we aim to shed light on the mechanisms through which AI systems emulate and possibly, enhance human capacities in artistic perception and creativity.
Macro-Queries: An Exploration into Guided Chart Generation from High Level Prompts
Lee, Christopher J., Tran, Giorgio, Tabalba, Roderick, Leigh, Jason, Longman, Ryan
This paper explores the intersection of data visualization and Large Language Models (LLMs). Driven by the need to make a broader range of data visualization types accessible for novice users, we present a guided LLM-based pipeline designed to transform data, guided by high-level user questions (referred to as macro-queries), into a diverse set of useful visualizations. This approach leverages various prompting techniques, fine-tuning inspired by Abela's Chart Taxonomy, and integrated SQL tool usage.
Controllable Text Generation for Large Language Models: A Survey
Liang, Xun, Wang, Hanyu, Wang, Yezhaohui, Song, Shichao, Yang, Jiawei, Niu, Simin, Hu, Jie, Liu, Dan, Yao, Shunyu, Xiong, Feiyu, Li, Zhiyu
In Natural Language Processing (NLP), Large Language Models (LLMs) have demonstrated high text generation quality. However, in real-world applications, LLMs must meet increasingly complex requirements. Beyond avoiding misleading or inappropriate content, LLMs are also expected to cater to specific user needs, such as imitating particular writing styles or generating text with poetic richness. These varied demands have driven the development of Controllable Text Generation (CTG) techniques, which ensure that outputs adhere to predefined control conditions--such as safety, sentiment, thematic consistency, and linguistic style--while maintaining high standards of helpfulness, fluency, and diversity. This paper systematically reviews the latest advancements in CTG for LLMs, offering a comprehensive definition of its core concepts and clarifying the requirements for control conditions and text quality. We categorize CTG tasks into two primary types: content control and attribute control. The key methods are discussed, including model retraining, fine-tuning, reinforcement learning, prompt engineering, latent space manipulation, and decoding-time intervention. We analyze each method's characteristics, advantages, and limitations, providing nuanced insights for achieving generation control. Additionally, we review CTG evaluation methods, summarize its applications across domains, and address key challenges in current research, including reduced fluency and practicality. We also propose several appeals, such as placing greater emphasis on real-world applications in future research. This paper aims to offer valuable guidance to researchers and developers in the field. Our reference list and Chinese version are open-sourced at https://github.com/IAAR-Shanghai/CTGSurvey.
Building and better understanding vision-language models: insights and future directions
Laurenรงon, Hugo, Marafioti, Andrรฉs, Sanh, Victor, Tronchon, Lรฉo
The field of vision-language models (VLMs), which take images and texts as inputs and output texts, is rapidly evolving and has yet to reach consensus on several key aspects of the development pipeline, including data, architecture, and training methods. This paper can be seen as a tutorial for building a VLM. We begin by providing a comprehensive overview of the current state-of-the-art approaches, highlighting the strengths and weaknesses of each, addressing the major challenges in the field, and suggesting promising research directions for underexplored areas. We then walk through the practical steps to build Idefics3-8B, a powerful VLM that significantly outperforms its predecessor Idefics2-8B, while being trained efficiently, exclusively on open datasets, and using a straightforward pipeline. These steps include the creation of Docmatix, a dataset for improving document understanding capabilities, which is 240 times larger than previously available datasets. We release the model along with the datasets created for its training.
Deal reached in feud between California news outlets and Google: 250 million to support journalism but no new law
California lawmakers intend to shelve legislation that would have required Google to pay news outlets for distributing their content, and in its place announced a new public-private partnership between the state and the tech giant that will fund programs to research artificial intelligence and bolster local journalism. The plan lays out a commitment of nearly 250 million over the next five years, with one-fourth of the money coming from state taxpayers and three-fourths of it coming from Google and possibly other private donors. The money will go toward two new initiatives administered by UC Berkeley's Graduate School of Journalism: a fund to distribute millions of dollars to California news outlets, and an "AI accelerator" to develop ways for journalists to use the powerful technology. "This agreement represents a major breakthrough in ensuring the survival of newsrooms and bolstering local journalism across California -- leveraging substantial tech industry resources without imposing new taxes on Californians," Gov. Gavin Newsom said in a statement. "The deal not only provides funding to support hundreds of new journalists, but helps rebuild a robust and dynamic California press corps for years to come, reinforcing the vital role of journalism in our democracy."
Writer calls for more working-class people in TV
Elsewhere in his speech, Graham championed public service broadcasters, which he said must not be taken for granted. Referring to critics of the BBC, he said: "Don't they realise that without the BBC, we lose our competitive advantage over the US markets? That not-for-profit means British stories, set in British communities, with British characters are protected by the licence fee, and may disappear without it?" Elsewhere, Graham suggested the new Labour government "should allow culture to play an active part in this promised national renewal โ not just kept at arms-length in its own silo on the peripheries of policy making, as it so often is". He added: "Creativity and arts subjects have been systematically stripped from the education system in England over the past 15 years. "A reduction of nearly half of all drama teachers, gone from those state schools since 2010.
Amazon back-to-school sale: 16 deals you can't miss
Save big on back to school and dorm room essentials on Amazon. During the back-to-school shopping season, you can find deep discounts on top brands on Amazon. Now is your chance to stock up on school supplies, backpacks and dorm room essentials for up to 50% off the list price. Get your back-to-school discounts delivered in time for the first day by signing up for a Prime membership. The benefits include fast, free delivery, access to invite-only deals and the option to Buy With Prime. Most purchases can be delivered to your door in 24 hours if you're an Amazon Prime member.
A Unified Framework for Continual Learning and Machine Unlearning
Chatterjee, Romit, Chundawat, Vikram, Tarun, Ayush, Mali, Ankur, Mandal, Murari
Continual learning and machine unlearning are crucial challenges in machine learning, typically addressed separately. Continual learning focuses on adapting to new knowledge while preserving past information, whereas unlearning involves selectively forgetting specific subsets of data. In this paper, we introduce a novel framework that jointly tackles both tasks by leveraging controlled knowledge distillation. Our approach enables efficient learning with minimal forgetting and effective targeted unlearning. By incorporating a fixed memory buffer, the system supports learning new concepts while retaining prior knowledge. The distillation process is carefully managed to ensure a balance between acquiring new information and forgetting specific data as needed. Experimental results on benchmark datasets show that our method matches or exceeds the performance of existing approaches in both continual learning and machine unlearning. This unified framework is the first to address both challenges simultaneously, paving the way for adaptable models capable of dynamic learning and forgetting while maintaining strong overall performance.
MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing
Zhou, Hao, Wang, Zhijun, Huang, Shujian, Huang, Xin, Han, Xue, Feng, Junlan, Deng, Chao, Luo, Weihua, Chen, Jiajun
Large Language Models (LLMs) are often English-centric due to the disproportionate distribution of languages in their pre-training data. Enhancing non-English language capabilities through post-pretraining often results in catastrophic forgetting of the ability of original languages. Previous methods either achieve good expansion with severe forgetting or slight forgetting with poor expansion, indicating the challenge of balancing language expansion while preventing forgetting. In this paper, we propose a method called MoE-LPR (Mixture-of-Experts with Language Priors Routing) to alleviate this problem. MoE-LPR employs a two-stage training approach to enhance the multilingual capability. First, the model is post-pretrained into a Mixture-of-Experts (MoE) architecture by upcycling, where all the original parameters are frozen and new experts are added. In this stage, we focus improving the ability on expanded languages, without using any original language data. Then, the model reviews the knowledge of the original languages with replay data amounting to less than 1% of post-pretraining, where we incorporate language priors routing to better recover the abilities of the original languages. Evaluations on multiple benchmarks show that MoE-LPR outperforms other post-pretraining methods. Freezing original parameters preserves original language knowledge while adding new experts preserves the learning ability. Reviewing with LPR enables effective utilization of multilingual knowledge within the parameters. Additionally, the MoE architecture maintains the same inference overhead while increasing total model parameters. Extensive experiments demonstrate MoE-LPR's effectiveness in improving expanded languages and preserving original language proficiency with superior scalability. Code and scripts are freely available at https://github.com/zjwang21/MoE-LPR.git.