Goto

Collaborating Authors

 Country




c46489a2d5a9a9ecfc53b17610926ddd-Supplemental.pdf

Neural Information Processing Systems

Since our manually collected test sets are rather small, we decided to avoid tuning hyperparameters on them as this would require holding out a non-trivial number of data points.



f14bc21be7eaeed046fed206a492e652-Supplemental.pdf

Neural Information Processing Systems

The major differences are as follows: 1) we use dropout and BN with weight normalization (WN) as a regularizer instead of the existing techniques such as spectral normalization (SN) and gradient penalties (GP). The BN is proven to function as a regularizer imposing the Lipschitz constraint [9], which has been achieved by SN and GP [3, 7]. Plus, the dropout and WN have been successfully adopted in the classifier-based model[12]. The learning rates of the discriminator and the generator are set according to two-timescale learning rate (TTUR) [4], which is adopted in Proj. SNGAN sets the learning rates ofthe discriminator and the generator as0.0004 and 0.0001, respectively,andtheyarefixedoverthecourse ofthetraining.




c460dc0f18fc309ac07306a4a55d2fd6-Paper.pdf

Neural Information Processing Systems

However,twokeydrawbacks of RPS-l2 are thatthey(i)leadtodisagreement between theoriginally trained networkandthe RPS-l2 regularized network modification and (ii) often yield a static ranking of training data fortest points inthesame class, independent ofthetest point being classified. Inspired by the RPS-l2 approach, we propose an alternative method based on a local Jacobian Taylor expansion (LJE).