Goto

Collaborating Authors

 Country



A.1 EquivalencebetweenPreConvSAandvanillaSA OurproposedPreConvSAisformulatedinEquation1. f0i =MLPs fli fl+1

Neural Information Processing Systems

Inspired by the depthwise separable convolution used in MobileNet [1], we extend this separable idea to graph convolution. We only perform MLPs on point features directly to learn thechannelcorrelation, andleverages theanisotropic reduction toaggregatethespatial correlation.








Exposing Attention Glitches with Flip-Flop Language Modeling

Neural Information Processing Systems

This simple generative task requires a model to copy binary symbols over long-range dependencies, ignoring the tokens in between. We find that Transformer FFLMs suffer from a long tail of sporadic reasoning errors, some of which we can eliminate using various regularization techniques.


Exposing Attention Glitches with Flip-Flop Language Modeling

Neural Information Processing Systems

This simple generative task requires a model to copy binary symbols over long-range dependencies, ignoring the tokens in between. We find that Transformer FFLMs suffer from a long tail of sporadic reasoning errors, some of which we can eliminate using various regularization techniques.