Goto

Collaborating Authors

 Statistical Learning


Label Noise SGD Provably Prefers Flat Global Minimizers

Neural Information Processing Systems

In overparametrized models, the noise in stochastic gradient descent (SGD) implicitly regularizes the optimization trajectory and determines which local minimum SGD converges to.


Appendix: A Lower Bound of Hash Codes ' Performance

Neural Information Processing Systems

All positives have ranks i placed on upper-right. We assume distances between query and any positive samples are different with each other. Figure 2: Mis-ranks marked on true positives and swaps that change ranks and mis-ranks. More generally, any swaps happen in a rank list would influence ranks and mis-ranks of involved positive samples. Eq. (3) is immediately obtained since B.1 Analysis on the Proposed Lower Bound B.1.1 Is the Introduced Lower Bound Tight?



SGD: The Role of Implicit Regularization, Batch-size and Multiple Epochs

Neural Information Processing Systems

Our main contributions are threefold: 1. We show that for any regularizer, there is an SCO problem for which Regularized Empirical Risk Minimzation fails to learn.


A Learning Algorithm Algorithm 1: Learning algorithm for Dr.k-NN Input: S

Neural Information Processing Systems

B.1 Proof of Theorem 1 The proof of Theorem 1 is based on the following two lemmas. Moreover, when there is a tie (i.e., the set Proof of Lemma 2. Recall that the Wasserstein metric of order 1 is defined as W ( P,P For the sake of completeness, we extend our algorithm to non-few-training-sample setting. The depth of the shaded area shows the level of samples entropy. The entropy of a sample is defined as follows. As a simple example, for Bernoulli random variable (which can represent, e.g., the outcome for flipping a coin with bias Now we use this entropy to define the "uncertainty" associated with each training points. Figure 6 reveals that the most informative samples usually lie in between categories.




Towards Understanding Why Lookahead Generalizes Better Than SGD and Beyond (Supplementary File) Pan Zhou Hanshu Y an

Neural Information Processing Systems

It is structured as follows. Then Appendix D gives the proofs of the main results in Sec. 4, including Finally, Appendix E provides the proofs of the results in Sec. 5, including Theorems 5 and 6 which analyze the optimization error, generalization error and excess risk error of the The main limitation of this work is that the analysis in this work cannot be applicable to general nonconvex problems. This is because as explained in Sec. But as shown in Sec. 3, to bound the excess risk error, one needs to first bound In this way, our analysis cannot be applicable to general nonconvex problems. Due to space limitation, we defer more experimental results and details to this appendix.



Supplementary Materials Rashomon Capacity: A Metric for Predictive Multiplicity in Classification

Neural Information Processing Systems

(since we pick the log base to be 2). We now prove the converse statements. Individual fairness aims to ensure that "similar individuals are treated similarly." Predictive multiplicity allows different predictions from competing classifiers for the samples. Notably, neural networks with very narrows or wide layers have better reproducibility in their decision regions. The fact that multiple classifiers may yield distinct predictions to a target a sample while having statistically identical average loss performance can also cause security issues in machine learning.