Goto

Collaborating Authors

 Instructional Material


ToMAP: Training Opponent-Aware LLM Persuaders with Theory of Mind

arXiv.org Artificial Intelligence

Large language models (LLMs) have shown promising potential in persuasion, but existing works on training LLM persuaders are still preliminary. Notably, while humans are skilled in modeling their opponent's thoughts and opinions proactively and dynamically, current LLMs struggle with such Theory of Mind (ToM) reasoning, resulting in limited diversity and opponent awareness. To address this limitation, we introduce Theory of Mind Augmented Persuader (ToMAP), a novel approach for building more flexible persuader agents by incorporating two theory of mind modules that enhance the persuader's awareness and analysis of the opponent's mental state. Specifically, we begin by prompting the persuader to consider possible objections to the target central claim, and then use a text encoder paired with a trained MLP classifier to predict the opponent's current stance on these counterclaims. Our carefully designed reinforcement learning schema enables the persuader learns how to analyze opponent-related information and utilize it to generate more effective arguments. Experiments show that the ToMAP persuader, while containing only 3B parameters, outperforms much larger baselines, like GPT-4o, with a relative gain of 39.4% across multiple persuadee models and diverse corpora. Notably, ToMAP exhibits complex reasoning chains and reduced repetition during training, which leads to more diverse and effective arguments. The opponent-aware feature of ToMAP also makes it suitable for long conversations and enables it to employ more logical and opponent-aware strategies. These results underscore our method's effectiveness and highlight its potential for developing more persuasive language agents. Code is available at: https://github.com/ulab-uiuc/ToMAP.


DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning

arXiv.org Artificial Intelligence

Sparse-reward reinforcement learning (RL) can model a wide range of highly complex tasks. Solving sparse-reward tasks is RL's core premise, requiring efficient exploration coupled with long-horizon credit assignment, and overcoming these challenges is key for building self-improving agents with superhuman ability. Prior work commonly explores with the objective of solving many sparse-reward tasks, making exploration of individual high-dimensional, long-horizon tasks intractable. We argue that solving such challenging tasks requires solving simpler tasks that are relevant to the target task, i.e., whose achieval will teach the agent skills required for solving the target task. We demonstrate that this sense of direction, necessary for effective exploration, can be extracted from existing RL algorithms, without leveraging any prior information. To this end, we propose a method for directed sparse-reward goal-conditioned very long-horizon RL (DISCOVER), which selects exploratory goals in the direction of the target task. We connect DISCOVER to principled exploration in bandits, formally bounding the time until the target task becomes achievable in terms of the agent's initial distance to the target, but independent of the volume of the space of all tasks. We then perform a thorough evaluation in high-dimensional environments. We find that the directed goal selection of DISCOVER solves exploration problems that are beyond the reach of prior state-of-the-art exploration methods in RL.


The Cultural Mapping and Pattern Analysis (CMAP) Visualization Toolkit: Open Source Text Analysis for Qualitative and Computational Social Science

arXiv.org Artificial Intelligence

The CMAP (Cultural Mapping and Pattern Analysis) visualization toolkit is an open-source suite for analyzing and visualizing text data--from qualitative fieldnotes and in-depth interview transcripts to historical documents and web-scraped data such as message board posts or blogs. The toolkit is designed for scholars integrating pattern analysis, data visualization, and explanation in qualitative and/or computational social science (CSS). Despite the existence of off-the-shelf commercial qualitative data analysis software, there remains a shortage of highly scalable open-source options capable of handling large datasets and supporting advanced statistical and language modeling. The foundation of the toolkit is a pragmatic approach that aligns research tools with social science project goals--empirical explanation, theory-guided measurement, comparative design, or evidence-based recommendations--guided by the principle that research paradigms and questions should determine methods. Consequently, the CMAP visualization toolkit offers a wide range of possibilities through the adjustment of a relatively small number of parameters and allows seamless integration with other Python tools.


Co-Designing Interdisciplinary Design Projects with AI

arXiv.org Artificial Intelligence

T his work has been submitted to the IEEE for possible publication. ORCID: 0000 -0003-2811-1194 Abstract --Creating interdisciplinary design projects is time-consuming and cognitively demanding for teachers, requiring curriculum alignment, cross -subject integration, and careful sequencing. This paper presents the Interdisciplinary Design Project Planner (IDPplanner), a GPT -based planning assistant grounded in Design Innovation principles, al ignment with Singapore secondary school's syllabuses, and 21st -century competencies. In a within -subject, counterbalanced workshop with 33 in -service teachers, participants produced two versions of the same project: manual and AI -assisted, followed by self - and peer-evaluations using a six -dimensional rubric. AI -assisted version received higher scores for Curriculum Alignment, Design Thinking Application, and Coherence & Flow, with a marginal advantage for Assessment Strategies. Teacher reflections indicated that AI -assisted planning improved structure, sequencing, and idea generation, while contextualization to local syllabuses, class profiles, and student needs remained teacher-led. Contributions include (1) a purpose-built planning tool that organizes ideas into a ten - component flow with ready-to -adapt prompts, templates, and assessment suggestions; (2) an empirical, rubric -based comparison of plan ning quality; and (3) evidence that AI can function as a pedagogical planning partner . Recommendations emphasize hybrid teacher-AI workflows to enhance curriculum alignment and reduce planning complexity, and design suggestions for developers to strengthen contextual customization, iterative design support, and l ocalized rubrics. Although instantiated with a Singapore -based curriculum, the planning flow and rubric are framework -agnostic and can be parameterized for other systems. Interdisciplinary learning approaches have gained prominence globally, particularly as countries prioritize 21st-century competencies (21CC) such as creativity, problem - solving, collaboration, and adaptive thinking.


FinFlowRL: An Imitation-Reinforcement Learning Framework for Adaptive Stochastic Control in Finance

arXiv.org Artificial Intelligence

Traditional stochastic control methods in finance struggle in real world markets due to their reliance on simplifying assumptions and stylized frameworks. Such methods typically perform well in specific, well defined environments but yield suboptimal results in changed, non stationary ones. We introduce FinFlowRL, a novel framework for financial optimal stochastic control. The framework pretrains an adaptive meta policy learning from multiple expert strategies, then finetunes through reinforcement learning in the noise space to optimize the generative process. By employing action chunking generating action sequences rather than single decisions, it addresses the non Markovian nature of markets. FinFlowRL consistently outperforms individually optimized experts across diverse market conditions.


Online Policy Learning via a Self-Normalized Maximal Inequality

arXiv.org Machine Learning

Adaptive experiments produce dependent data that break i.i.d. assumptions that underlie classical concentration bounds and invalidate standard learning guarantees. In this paper, we develop a self-normalized maximal inequality for martingale empirical processes. Building on this, we first propose an adaptive sample-variance penalization procedure which balances empirical loss and sample variance, valid for general dependent data. Next, this allows us to derive a new variance-regularized pessimistic off-policy learning objective, for which we establish excess-risk guarantees. Subsequently, we show that, when combined with sequential updates and under standard complexity and margin conditions, the resulting estimator achieves fast convergence rates in both parametric and nonparametric regimes, improving over the usual $1/\sqrt{n}$ baseline. We complement our theoretical findings with numerical simulations that illustrate the practical gains of our approach.


Particle Dynamics for Latent-Variable Energy-Based Models

arXiv.org Machine Learning

Latent-variable energy-based models (LV-EBMs) assign a single normalized energy to joint pairs of observed data and latent variables, offering expressive generative modeling while capturing hidden structure. We recast maximum-likelihood training as a saddle problem over distributions on the latent and joint manifolds and view the inner updates as coupled Wasserstein gradient flows. The resulting algorithm alternates overdamped Langevin updates for a joint negative pool and for conditional latent particles with stochastic parameter ascent, requiring no discriminator or auxiliary networks. We prove existence and convergence under standard smoothness and dissi-pativity assumptions, with decay rates in KL divergence and Wasserstein-2 distance. The saddle-point view further yields an ELBO strictly tighter than bounds obtained with restricted amortized posteriors. Our method is evaluated on numerical approximations of physical systems and performs competitively against comparable approaches.


Nonlinear Dimensionality Reduction Techniques for Bayesian Optimization

arXiv.org Artificial Intelligence

Bayesian optimisation (BO) is a standard approach for sample-efficient global optimisation of expensive black-box functions, yet its scalability to high dimensions remains challenging. Here, we investigate nonlinear dimensionality reduction techniques that reduce the problem to a sequence of low-dimensional Latent-Space BO (LSBO). While early LSBO methods used (linear) random projections (Wang et al., 2013), building on Grosnit et al. (2021), we employ Variational Autoencoders (VAEs) for LSBO, focusing on deep metric loss for structured latent manifolds and VAE retraining to adapt the encoder-decoder to newly sampled regions. We propose some changes in their implementation, originally designed for tasks such as molecule generation, and reformulate the algorithm for broader optimisation purposes. We then couple LSBO with Sequential Domain Reduction (SDR) directly in the latent space (SDR-LSBO), yielding an algorithm that narrows the latent search domains as evidence accumulates. Implemented in a GPU-accelerated BoTorch stack with Matern-5/2 Gaussian process surrogates, our numerical results show improved optimisation quality across benchmark tasks and that structured latent manifolds improve BO performance. Additionally, we compare random embeddings and VAEs as two mechanisms for dimensionality reduction, showing that the latter outperforms the former. To the best of our knowledge, this is the first study to combine SDR with VAE-based LSBO, and our analysis clarifies design choices for metric shaping and retraining that are critical for scalable latent space BO. For reproducibility, our source code is available at https://github.com/L-Lok/Nonlinear-Dimensionality-Reduction-Techniques-for-Bayesian-Optimization.git.


4 ways to fix 'tech neck,' according to a physical therapist

Popular Science

Strengthening can help if you're staring at your phone too much. You don't need a ton of equipment to fix your neck. Breakthroughs, discoveries, and DIY tips sent every weekday. If you're here seeking relief from tech neck, or the forward head posture associated with the use of personal devices, we've got good and bad news. The good news is you've come to the right place; the bad news is you're probably contributing to it right now.


Inside San Francisco's new AI school: is this the future of US education?

The Guardian

Experts have raised questions about whether an app-based curriculum can serve all learners equally. Experts have raised questions about whether an app-based curriculum can serve all learners equally. Inside San Francisco's new AI school: is this the future of US education? In the world's tech innovation epicenter, an "AI-powered" private school has made headlines for unabashedly embracing the technology. Alpha School San Francisco, which opened its doors to K-8 students this fall, is the newest outpost of a network of 14 nationwide private schools.