Goto

Collaborating Authors

 Technology


ImprovedFine-TuningbyBetterLeveraging Pre-TrainingData

Neural Information Processing Systems

As a dominant paradigm, fine-tuning a pre-trained model on the target data is widely used in many deep learning applications, especially for small data sets.



AppendixofSynergy-of-experts 1 TheoreticalProofs

Neural Information Processing Systems

From Figure 1(a), learning multiple linear sub-models and averaging the predictions (ensemble) is still a linear model, so it cannot tackleXOR problem. We compare the training cost of all methods from the two aspects;1). Thesub-model training enables themost adversarial attacks ofsub-models could be successfully defended. In particular, we train two kinds of models to defend against the attacks: 1). FromFigure2(a)and2(b),when0.01 ฯต 0.04, SoE without the collaboration training achieves a similar robustness compared with SoE.








ErrorCompensatedDistributedSGD canbeAccelerated

Neural Information Processing Systems

In this work, we show for the first time that error compensated gradient compression methods can be accelerated. In particular, we propose and study the error compensated loopless Katyusha method, and establish an accelerated linear convergence rate under standard assumptions.