Goto

Collaborating Authors

 Large Language Model






BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts

Neural Information Processing Systems

The Mixture of Experts (MoE) framework has become a popular architecture for large language models due to its superior performance over dense models. However, training MoEs from scratch in a large-scale regime is prohibitively expensive.



Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level Signature

Neural Information Processing Systems

Text watermarks for large language models (LLMs) have been commonly used to identify the origins of machine-generated content, which is promising for assessing liability when combating deepfake or harmful content.




WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models

Neural Information Processing Systems

Despite growing interest, much of the existing research has focused on varied unlearning method designs to boost effectiveness and efficiency. However, the inherent relationship between model weights and LLM unlearning has not been extensively examined.