Goto

Collaborating Authors

 Asia


Representation Noising: A Defence Mechanism Against Harmful Finetuning

Neural Information Processing Systems

Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release of weights, weight stealing and fine-tuning APIs make closed models vulnerable to harmful fine-tuning attacks (HFAs).



SupplementaryMaterials

Neural Information Processing Systems

Efficiency: The overall reward can be allocated to all players in the game,i.e. This section provides more details about multi-order interactions [8] in Section 3.3 of the paper. The multi-order interaction satisfies axioms oflinearity, nullity, commutativity, symmetry, and efficiency[8],asfollows. This study was done under the supervision of Dr. Quanshi Zhang. This section provides more details about the use of the ShapeNet part dataset in the paper.