Goto

Collaborating Authors

 Large Language Model








HuRef: HUman-REadable Fingerprint for Large Language Models

Neural Information Processing Systems

However, identifying the original base model of an LLM is challenging due to potential parameter alterations. In this study, we introduce HuRef, a human-readable fingerprint for LLMs that uniquely identifies the base model without interfering with training or exposing model parameters to the public.



Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models

Neural Information Processing Systems

Since the rapid development of Large Language Models (LLMs) has achieved remarkable success, understanding and rectifying their internal complex mechanisms has become an urgent issue. Recent research has attempted to interpret their behaviors through the lens of inner representation. However, developing practical and efficient methods for applying these representations for general and flexible model editing remains challenging.