Whilemuchwork has focused on different weight pruning criteria, the overallsparsifiabilityofthe network, i.e., its capacity to be pruned without quality loss, has often been overlooked.
Classical methods, such as weighted cross-entropy, fail when training deep nets to the terminal phase of training (TPT), that is training beyond zero training error.
Weintroduce Unbalanced SobolevDescent (USD), aparticle descent algorithm for transporting a high dimensional source distribution to a target distribution that does not necessarily have the same mass.