Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
Vasilev, Stefan, Herold, Christian, Liao, Baohao, Hashemi, Seyyed Hadi, Khadivi, Shahram, Monz, Christof
–arXiv.org Artificial Intelligence
This paper introduces Unilogit, a novel self-distillation method for machine unlearning in Large Language Models. Unilogit addresses the challenge of selectively forgetting specific information while maintaining overall model utility, a critical task in compliance with data privacy regulations like GDPR. Unlike prior methods that rely on static hyperparameters or starting model outputs, Unilogit dynamically adjusts target logits to achieve a uniform probability for the target token, leveraging the current model's outputs for more accurate self-distillation targets. This approach not only eliminates the need for additional hy-perparameters but also enhances the model's ability to approximate the golden targets. Extensive experiments on public benchmarks and an in-house e-commerce dataset demonstrate Unilogit's superior performance in balancing forget and retain objectives, outperforming state-of-the-art methods such as NPO and Un-DIAL. Our analysis further reveals Unilogit's robustness across various scenarios, highlighting its practical applicability and effectiveness in achieving efficacious machine unlearning.
arXiv.org Artificial Intelligence
May-12-2025
- Country:
- Asia > Middle East > UAE (0.28)
- Genre:
- Research Report
- New Finding (0.93)
- Promising Solution (0.89)
- Research Report
- Industry:
- Information Technology
- Security & Privacy (1.00)
- Services > e-Commerce Services (0.35)
- Information Technology
- Technology: