Goto

Collaborating Authors

 Deep Learning


Supplement to: Embedding Principle of Loss Landscape of Deep Neural Networks

Neural Information Processing Systems

However, this transform does not inform about the degeneracy of critical points/manifolds. Clearly, this transform is also a critical transform. For the 1D fitting experiments (Figs. 1, 3(a), 4), we use tanh as the activation function, mean squared We use the full-batch gradient descent with learning rate 0.005 to We use the default Adam optimizer of full batch with learning rate 0.02 to train for We also use the default Adam optimizer of full batch with learning rate 0.00003 Their output functions are shown in the figure. Remark that, although Figs. 1 and 5 are case studies each based on a random trial, similar phenomenon Do the main claims made in the abstract and introduction accurately reflect the paper's Did you state the full set of assumptions of all theoretical results? Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Y es] In the Did you specify all the training details (e.g., data splits, hyperparameters, how they Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)?


Embedding Principle of Loss Landscape of Deep Neural Networks

Neural Information Processing Systems

Understanding the structure of loss landscape of deep neural networks (DNNs) is obviously important. In this work, we prove an embedding principle that the loss landscape of a DNN "contains" all the critical points of all the narrower DNNs.


A Observations in Local Memory Similarity

Neural Information Processing Systems

We observed local memory's similarity through Q-Q (quantile-quantile) plots as shown in Figure In Figure A1(a), the linearity of the points in Q-Q plot suggests that the worker 1's local This is consistent to our observations in pairwise cosine distance shown in Figure 2(a). This indicates that we can possibly use local worker's top-k One variant of Y oung's inequality is k x + y k A.1 global minimum of f ( x) 2, The quadrilateral identity is h x, y i = 1 2 k x k We provided the following table to explain section 3's main results and connected them to other parts of paper. Our theorem 1 shows this; indicates its applicability in distributed training. Lemma1: contraction property Lemma2: contraction in distributed setting Theorem1: ScaleCom's convergence rate same as SGD ( 1 / p T) Intuition Higher correlation between workers brings CL T - k closer to true top-k Require positive correlation between workers in distr. Fig.2 and 3 show high correlation so our contraction is close to true top-k Fig.2 and 3 show positive correlation between workers Table 1,2 (Fig4,5) verified ScaleCom's convergence same as baseline Each node is equipped with 2 IBM Power 9 processors clocked at 3.15 GHz.


ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed Training Chia-Y u Chen

Neural Information Processing Systems

Large-scale distributed training of Deep Neural Networks (DNNs) on state-of-the-art platforms is expected to be severely communication constrained. To overcome this limitation, numerous gradient compression techniques have been proposed and have demonstrated high compression ratios. However, most existing methods do not scale well to large scale distributed systems (due to gradient build-up) and/or fail to evaluate model fidelity (test accuracy) on large datasets.





Physics-Integrated Variational Autoencoders for Robust and Interpretable Generative Modeling

Neural Information Processing Systems

A technical challenge in deep gray-box modeling is to ensure an appropriate use of physics models. A careless design of models and learning can lead to an erratic behavior of the components meant to represent physics (e.g., with erroneous estimation of physics parameters), and eventually, the overall


Online Meta-Learning via Learning with Layer-Distributed Memory

Neural Information Processing Systems

We demonstrate that efficient meta-learning can be achieved via end-to-end training of deep neural networks with memory distributed across layers. The persistent state of this memory assumes the entire burden of guiding task adaptation. Moreover, its distributed nature is instrumental in orchestrating adaptation.


A Appendix

Neural Information Processing Systems

Same as other neural networks, Transformer-based models use floating-point arithmetic, however cryptographic protocols operate on integers. As shown in Section 4.1, the matrix multiplication Note that this is quite fast due to the small search range. As shown in Section 3.1 of [ The correctness of Theorem A.2 is directly derived from Theorem 4.1. Our security proof follows the simulation paradigm defined in Section A.1.2. In our setting, m is always less than N .