Goto

Collaborating Authors

 Media


Aligning Large Language Models with Representation Editing: A Control Perspective

Neural Information Processing Systems

Aligning large language models (LLMs) with human objectives is crucial for real-world applications. However, fine-tuning LLMs for alignment often suffers from unstable training and requires substantial computing resources. Test-time alignment techniques, such as prompting and guided decoding, do not modify the underlying model, and their performance remains dependent on the original model's capabilities. To address these challenges, we propose aligning LLMs through representation editing. The core of our method is to view a pre-trained autoregressive LLM as a discrete-time stochastic dynamical system.







Using a swearword in your Google search can stop the AI answer. But should you?

The Guardian

Using a swearword in your Google search can stop the AI answer. Artificial intelligence is more than Trump deepfakes of Tilly the actor. It's used in smartphones, customer service, healthcare - even legal cases. Is it possible to avoid? Using a swearword in your Google search can stop that annoying AI overview from popping up.