Goto

Collaborating Authors

 Deep Learning



e7d019329e662fe4685be505befca3bb-Paper-Conference.pdf

Neural Information Processing Systems

Inductive biases encoding known data symmetries are key to make deep learning models generalize in high-dimensional settings such as computer vision, speech processing and computational neuroscience, just to name a few.


VeLoRA: MemoryEfficientTrainingusing Rank-1Sub-TokenProjections

Neural Information Processing Systems

Using a single projection vector, we then project these individual sub-tokens onto a one-dimensional subspace. Importantly, we notice that we can initialize this projection vector cheaply using first-order batch statistics andthen keepitfixedthroughout training. Wethen reconstruct the original tokens using the same vector during the backward pass.