Goto

Collaborating Authors

 Deep Learning






Pay attention to your loss: understanding misconceptions about 1-Lipschitz neural networks

Neural Information Processing Systems

However they remain commonly considered as less accurate, and their properties in learning are still not fully understood. In this paper we clarify the matter: when it comes to classification 1-Lipschitz neural networks enjoy several advantages over their unconstrained counterpart.






Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models Xiuying Wei

Neural Information Processing Systems

Therefore, transformer quantization attracts wide research interest. Recent work recognizes that structured outliers are the critical bottleneck for quantization performance. However, their proposed methods increase the computation overhead and still leave the outliers there. To fundamentally address this problem, this paper delves into the inherent inducement and importance of the outliers.