Goto

Collaborating Authors

 Deep Learning



Towards Practical Control of Singular Values of Convolutional Layers

Neural Information Processing Systems

In general, convolutional neural networks (CNNs) are easy to train, but their essential properties, such as generalization error and adversarial robustness, are hard to control. Recent research demonstrated that singular values of con-volutional layers significantly affect such elusive properties and offered several methods for controlling them. Nevertheless, these methods present an intractable computational challenge or resort to coarse approximations. In this paper, we offer a principled approach to alleviating constraints of the prior art at the expense of an insignificant reduction in layer expressivity. Our method is based on the tensor-train decomposition; it retains control over the actual singular values of convolutional mappings while providing structurally sparse and hardware-friendly representation. We demonstrate the improved properties of modern CNNs with our method and analyze its impact on the model performance, calibration, and adversarial robustness.


Direct Feedback Alignment Scales to Modern Deep Learning Tasks and Architectures

Neural Information Processing Systems

Despite being the workhorse of deep learning, the backpropagation algorithm is no panacea. It enforces sequential layer updates, thus preventing efficient paral-lelization of the training process.





Convergence beyond the over-parameterized regime using Rayleigh quotients

Neural Information Processing Systems

In this paper, we present a new strategy to prove the convergence of deep learning architectures to a zero training (or even testing) loss by gradient flow. Our analysis is centered on the notion of Rayleigh quotients in order to prove Kurdyka-ลojasiewicz inequalities for a broader set of neural network architectures and loss functions.