Large Language Model Unlearning
–Neural Information Processing Systems
We study how to perform unlearning, i.e. forgetting undesirable (mis)behaviors, on large language models (LLMs). Unlearning, as an alignment technique, has three advantages. To the best of our knowledge, our work is among the first to explore LLM unlearning. We are also among the first to formulate the settings, goals, and evaluations in LLM unlearning. Despite only having negative samples, our ablation study shows that unlearning can still achieve better alignment performance than RLHF with just 2% of its computational time.
Neural Information Processing Systems
May-27-2025, 14:57:06 GMT
- Technology: