Goto

Collaborating Authors

 Statistical Learning








Appendices A HSIC estimation in the self-supervised setting

Neural Information Processing Systems

Estimators of HSIC typically assume i.i.d. A.3 Estimator of HSIC(Z, Z) Before discussing estimators of HSIC(Z,Z), note that it takes the following form: HSIC(Z,Z) = E null k (Z,Z Finally, note that even if null HSIC(Z,Z) is unbiased, its square root is not. B.1 InfoNCE connection To establish the connection with InfoNCE, define it in terms of expectations: L In the small variance regime, InfoNCE also bounds an HSIC-based loss. Both roots are real, as ฮฑ 1 /4. Theorem B.1 works for any bounded kernel, because In Section 3.2, we make the assumption that the features are centered and argue that the assumption is valid for BYOL.




Appendix A Gradient Descent and Neural Tangent Kernel Gradient Descent Since we consider the square loss and `

Neural Information Processing Systems

We provide here a brief overview of reproducing kernel Hilbert space (RKHS). More details can be found in Appendix G.2. In this work, we impose the following assumptions. Remark 5. Assumption D.3 can be replaced by an alternative assumption, that is, Assumption D.1 is related to the neural network and GD training, where similar settings have been Assumption D.2 imposes conditions on the underlying true conditional probability in the non-separable case. This assumption basically requires that the conditional probability is within the function class generated by the GD-trained neural networks we consider (thus can be calibrated).