Large Language Model
Scaling laws for activation steering with Llama 2 models and refusal mechanisms
Ali, Sheikh Abdur Raheem, Xu, Justin, Yang, Ivory, Li, Jasmine Xinze, Arslan, Ayse, Benham, Clark
As large language models (LLMs) evolve in complexity and capability, the efficacy of less widely deployed alignment techniques are uncertain. Building on previous work on activation steering and contrastive activation addition (CAA), this paper explores the effectiveness of CAA with model scale using the family of Llama 2 models (7B, 13B, and 70B). CAA works by finding desirable 'directions' in the model's residual stream vector space using contrastive pairs (for example, hate to love) and adding this direction to the residual stream during the forward pass. It directly manipulates the residual stream and aims to extract features from language models to better control their outputs. Using answer matching questions centered around the refusal behavior, we found that 1) CAA is most effective when applied at early-mid layers. 2) The effectiveness of CAA diminishes with model size. 3) Negative steering has more pronounced effects than positive steering across all model sizes.
Auto-Formulating Dynamic Programming Problems with Large Language Models
Zhou, Chenyu, Yang, Jingyuan, Xin, Linwei, Chen, Yitian, He, Ziyan, Ge, Dongdong
Automating the formulation of decision-making problems represents a major step toward fully autonomous decision-support systems. Traditionally, solving such problems involves two sequential stages: first, translating real-world scenarios into well-defined mathematical models-an essential skill emphasized in operations research education-and second, applying computational tools to find optimal or near-optimal solutions. While substantial research in recent decades has primarily focused on the second stage-enhancing algorithms and improving solver efficiency-advancements span a wide range, from foundational developments such as reinforcement learning (RL) frameworks (e.g., Sutton and Barto 2018) and approximate dynamic programming techniques (e.g., Powell 2011), to powerful solvers like COPT, CPLEX, and Gurobi. Such innovations coupled with increasing computational power have led to high-impact real-world applications, exemplified by AlphaGo, which leveraged deep learning and RL to solve complex, large-scale decision-making problems (Silver et al. 2016). That said, while these advancements have shifted many computational tasks to automated software, the initial problem formulation step has largely remained manual and dependent on expert knowledge. The recent rapid progress in large language models (LLMs) provides a promising opportunity to automate this crucial first step. LLMs excel in natural language processing and have demonstrated significant potential for effectively automating the formulation of mathematical models directly from plain English descriptions. Leveraging LLMs can substantially reduce the human expertise required, simplify the problem formulation process, and make advanced optimization methods accessible to a broader audience. Among various optimization problems, dynamic programming (DP) represents a particularly important yet challenging category for formulation automation.
ExpliCIT-QA: Explainable Code-Based Image Table Question Answering
Lagos, Maximiliano Hormazรกbal, Sรกez, รlvaro Bueno, Doval, Pedro Alonso, Vesteiro, Jorge Alcalde, Cerezo-Costas, Hรฉctor
We present ExpliCIT-QA, a system that extends our previous MRT approach for tabular question answering into a multimodal pipeline capable of handling complex table images and providing explainable answers. ExpliCIT-QA follows a modular design, consisting of: (1) Multimodal Table Understanding, which uses a Chain-of-Thought approach to extract and transform content from table images; (2) Language-based Reasoning, where a step-by-step explanation in natural language is generated to solve the problem; (3) Automatic Code Generation, where Python/Pandas scripts are created based on the reasoning steps, with feedback for handling errors; (4) Code Execution to compute the final answer; and (5) Natural Language Explanation that describes how the answer was computed. The system is built for transparency and auditability: all intermediate outputs, parsed tables, reasoning steps, generated code, and final answers are available for inspection. This strategy works towards closing the explainability gap in end-to-end TableVQA systems. We evaluated ExpliCIT-QA on the TableVQA-Bench benchmark, comparing it with existing baselines. We demonstrated improvements in interpretability and transparency, which open the door for applications in sensitive domains like finance and healthcare where auditing results are critical.
General Modular Harness for LLM Agents in Multi-Turn Gaming Environments
Zhang, Yuxuan, Yu, Haoyang, Hu, Lanxiang, Jin, Haojian, Zhang, Hao
We introduce a modular harness design for LLM agents that composes of perception, memory, and reasoning components, enabling a single LLM or VLM backbone to tackle a wide spectrum of multi turn gaming environments without domain-specific engineering. Using classic and modern game suites as low-barrier, high-diversity testbeds, our framework provides a unified workflow for analyzing how each module affects performance across dynamic interactive settings. Extensive experiments demonstrate that the harness lifts gameplay performance consistently over un-harnessed baselines and reveals distinct contribution patterns, for example, memory dominates in long-horizon puzzles while perception is critical in vision noisy arcades. These findings highlight the effectiveness of our modular harness design in advancing general-purpose agent, given the familiarity and ubiquity of games in everyday human experience.
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
Dombrowski, Ann-Kathrin, Bowen, Dillon, Gleave, Adam, Cundy, Chris
Open-weight large language models (LLMs) unlock huge benefits in innovation, personalization, privacy, and democratization. However, their core advantage - modifiability - opens the door to systemic risks: bad actors can trivially subvert current safeguards, turning beneficial models into tools for harm. This leads to a 'safety gap': the difference in dangerous capabilities between a model with intact safeguards and one that has been stripped of those safeguards. We open-source a toolkit to estimate the safety gap for state-of-the-art open-weight models. As a case study, we evaluate biochemical and cyber capabilities, refusal rates, and generation quality of models from two families (Llama-3 and Qwen-2.5) across a range of parameter scales (0.5B to 405B) using different safeguard removal techniques. Our experiments reveal that the safety gap widens as model scale increases and effective dangerous capabilities grow substantially when safeguards are removed. We hope that the Safety Gap Toolkit (https://github.com/AlignmentResearch/safety-gap) will serve as an evaluation framework for common open-source models and as a motivation for developing and testing tamper-resistant safeguards. We welcome contributions to the toolkit from the community.
Reasoning Strategies in Large Language Models: Can They Follow, Prefer, and Optimize?
Zhang, Yanjian, Wisniewski, Guillaume, Tomeh, Nadi, Charnois, Thierry
Human reasoning involves different strategies, each suited to specific problems. Prior work shows that large language model (LLMs) tend to favor a single reasoning strategy, potentially limiting their effectiveness in diverse reasoning challenges. In this work, we investigate whether prompting can control LLMs reasoning strategies and assess its impact on logical problem-solving. While our experiments show that no single strategy consistently improves accuracy, performance could be enhanced if models could adaptively choose the optimal strategy. We propose methods to guide LLMs in strategy selection, highlighting new ways to refine their reasoning abilities.
Automated Novelty Evaluation of Academic Paper: A Collaborative Approach Integrating Human and Large Language Model Knowledge
Wu, Wenqing, Zhang, Chengzhi, Zhao, Yi
Novelty is a crucial criterion in the peer review process for evaluating academic papers. Traditionally, it's judged by experts or measure by unique reference combinations. Both methods have limitations: experts have limited knowledge, and the effectiveness of the combination method is uncertain. Moreover, it's unclear if unique citations truly measure novelty. The large language model (LLM) possesses a wealth of knowledge, while human experts possess judgment abilities that the LLM does not possess. Therefore, our research integrates the knowledge and abilities of LLM and human experts to address the limitations of novelty assessment. One of the most common types of novelty in academic papers is the introduction of new methods. In this paper, we propose leveraging human knowledge and LLM to assist pretrained language models (PLMs, e.g. BERT etc.) in predicting the method novelty of papers. Specifically, we extract sentences related to the novelty of the academic paper from peer review reports and use LLM to summarize the methodology section of the academic paper, which are then used to fine-tune PLMs. In addition, we have designed a text-guided fusion module with novel Sparse-Attention to better integrate human and LLM knowledge. We compared the method we proposed with a large number of baselines. Extensive experiments demonstrate that our method achieves superior performance.
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
Liu, Ziru, Gong, Cheng, Fu, Xinyu, Liu, Yaofang, Chen, Ran, Hu, Shoubo, Zhang, Suiyun, Liu, Rui, Zhang, Qingfu, Tu, Dandan
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a powerful paradigm for facilitating the self-improvement of large language models (LLMs), particularly in the domain of complex reasoning tasks. However, prevailing on-policy RL methods often contend with significant training instability and inefficiency. This is primarily due to a capacity-difficulty mismatch, where the complexity of training data frequently outpaces the model's current capabilities, leading to critically sparse reward signals and stalled learning progress. This challenge is particularly acute for smaller, more resource-efficient LLMs. To overcome this, we introduce the Guided Hybrid Policy Optimization (GHPO), a novel difficulty-aware reinforcement learning framework. GHPO dynamically calibrates task difficulty by employing adaptive prompt refinement to provide targeted guidance. This unique approach adaptively balances direct imitation learning for problems currently beyond the model's reach with exploration-based reinforcement learning for more manageable tasks, effectively creating a smooth and optimized learning curriculum. Extensive experiments demonstrate that GHPO achieves an average performance gain of approximately 5% across six challenging mathematics benchmarks, consistently outperforming strong on-policy reinforcement learning and curriculum learning baselines. Further analysis confirms that our framework significantly enhances both training stability and final reasoning performance, thus offering a scalable and efficient solution for developing powerful and robust reasoning models.
The Challenge of Teaching Reasoning to LLMs Without RL or Distillation
Du, Wei, Kisacanin, Branislav, Armstrong, George, Toshniwal, Shubham, Moshkov, Ivan, Ayrapetyan, Alexan, Mahdavi, Sadegh, Zhao, Dan, Diao, Shizhe, Masulovic, Dragan, Stanean, Marius, Avadhanam, Advaith, Wang, Max, Dutta, Ashmit, Govil, Shitij, Yanamandara, Sri, Tandon, Mihir, Ananthakrishnan, Sriram, Rathi, Vedant, Zhang, David, Kang, Joonseok, Luo, Leon, Andreescu, Titu, Ginsburg, Boris, Gitman, Igor
Reasoning-capable language models achieve state-of-the-art performance in diverse complex tasks by generating long, explicit Chain-of-Thought (CoT) traces. While recent works show that base models can acquire such reasoning traces via reinforcement learning or distillation from stronger models like DeepSeek-R1, previous works demonstrate that even short CoT prompting without fine-tuning is able to improve reasoning. We ask whether long CoT can be induced in a base model using only prompting or minimal tuning. Using just 20 long CoT examples from the reasoning model \texttt{QwQ-32B-Preview}, we lightly fine-tune the base model \texttt{Qwen2.5-32B}. The resulting model outperforms the much larger \texttt{Qwen2.5-Math-72B-Instruct}, showing that a handful of high-quality examples can unlock strong reasoning capabilities. We further explore using CoT data from non-reasoning models and human annotators, enhanced with prompt engineering, multi-pass editing, and structural guidance. However, neither matches the performance of reasoning model traces, suggesting that certain latent qualities of expert CoT are difficult to replicate. We analyze key properties of reasoning data, such as problem difficulty, diversity, and answer length, that influence reasoning distillation. While challenges remain, we are optimistic that carefully curated human-written CoT, even in small quantities, can activate reasoning behaviors in base models. We release our human-authored dataset across refinement stages and invite further investigation into what makes small-scale reasoning supervision so effective.
Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
Li, Yangning, Zhang, Weizhi, Yang, Yuyao, Huang, Wei-Chieh, Wu, Yaozu, Luo, Junyu, Bei, Yuanchen, Zou, Henry Peng, Luo, Xiao, Zhao, Yusheng, Chan, Chunkit, Chen, Yankai, Deng, Zhongfen, Li, Yinghui, Zheng, Hai-Tao, Li, Dongyuan, Jiang, Renhe, Zhang, Ming, Song, Yangqiu, Yu, Philip S.
Retrieval-Augmented Generation (RAG) lifts the factuality of Large Language Models (LLMs) by injecting external knowledge, yet it falls short on problems that demand multi-step inference; conversely, purely reasoning-oriented approaches often hallucinate or mis-ground facts. This survey synthesizes both strands under a unified reasoning-retrieval perspective. We first map how advanced reasoning optimizes each stage of RAG (Reasoning-Enhanced RAG). Then, we show how retrieved knowledge of different type supply missing premises and expand context for complex inference (RAG-Enhanced Reasoning). Finally, we spotlight emerging Synergized RAG-Reasoning frameworks, where (agentic) LLMs iteratively interleave search and reasoning to achieve state-of-the-art performance across knowledge-intensive benchmarks. We categorize methods, datasets, and open challenges, and outline research avenues toward deeper RAG-Reasoning systems that are more effective, multimodally-adaptive, trustworthy, and human-centric. The collection is available at https://github.com/DavidZWZ/Awesome-RAG-Reasoning.