UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
Lu, Huimin, Isonuma, Masaru, Mori, Junichiro, Sakata, Ichiro
–arXiv.org Artificial Intelligence
Previous detoxification methods are typically model-specific, addressing only individual models or model families, and require careful hyperparameter tuning due to the trade-off between detoxification efficacy and language modeling performance. Specifically, we propose a novel and efficient dataset distillation technique for detoxification using contrastive decoding. This approach distills detoxifying representations in the form of synthetic text data, enabling universal detoxification of any LLM through fine-tuning with the distilled text. Our experiments demonstrate that the detoxifying text distilled from GPT -2 can effectively detoxify larger models, including OPT, Falcon, and LLaMA-2. Additionally, analysis of the detoxifying text reveals a reduction in politically biased content, providing insights into the attributes necessary for effective detoxification of LLMs. Our codes are available at https://github.com/EminLU/UniDetox. Fascinated by the remarkable capabilities of Large Language Models (LLMs), numerous researchers and developers are dedicating their efforts to building new models. Today, many off-the-shelf pre-trained LLMs are publicly available (Radford et al., 2019; Zhang et al., 2022; Almazrouei et al., 2023; Touvron et al., 2023), and practitioners employ them in a wide range of applications. While this trend is expected to drive innovation across various fields, it simultaneously raises significant concerns regarding the unintended harmful behaviors exhibited by LLMs. LLMs, developed through pre-training on a large-scale corpus, often unintentionally acquire toxic content present in their training datasets (Gehman et al., 2020; Webster et al., 2020; Nozza et al., 2021). Due to these concerns, there have been efforts to introduce comprehensive regulations to mitigate the toxicity of LLMs; however, there is currently no standardized approach capable of consistently removing toxic content across diverse models. By developing a universal detoxification approach, we can form the basis for broadly applicable regulations and ensure consistent toxicity mitigation across a wide variety of LLMs. While numerous studies have explored the detoxification of LLMs, there is currently no post-hoc approach that can be seamlessly applied across models with varying architectures, sizes, or tok-enizers. Existing post-hoc detoxification strategies include decoding-time control (Liu et al., 2021; Zhang & Wan, 2023), word embedding/logits modification (Gehman et al., 2020; Han et al., 2024), and model editing (Ilharco et al., 2023; Wang et al., 2024). Crucially, this equilibrium point varies across models, necessitating individual hyperparameter optimization for each model, as we will thoroughly investigate in our experiments.
arXiv.org Artificial Intelligence
Apr-30-2025
- Country:
- North America > United States
- California (0.46)
- Missouri > Jackson County
- Kansas City (0.14)
- North America > United States
- Genre:
- Research Report > New Finding (1.00)
- Industry:
- Law Enforcement & Public Safety > Crime Prevention & Enforcement (1.00)
- Transportation (0.68)
- Government (0.67)
- Leisure & Entertainment > Sports
- Soccer (0.46)
- Technology: