Goto

Collaborating Authors

 Technology


I brought my husband back for his funeral as a hologram

BBC News

When Pam Cronrath's husband Bill died last year, after nearly 60 years of marriage, she knew what she wanted to do, but not exactly how. I promised him a super wake, she told the BBC. What she didn't expect was that keeping the promise would lead her into the world of holograms, technology more commonly associated with celebrities than memorial services in rural America. A self-confessed tech enthusiast, she says her outlook was shaped by a career that stretched back to the early days of the internet. Several years ago, while speaking at a medical conference, she watched a doctor appear as a full-body hologram broadcast live across the United States.




5446f217e9504bc593ad9dcf2ec88dda-Supplemental.pdf

Neural Information Processing Systems

Python notebooks producing the figures of this paper are available at https://github.com/ Let F be the joint distribution on RTD obtained by first assigning a random vector =( 1| | T) via F and then applying the map to each of the components t, and let F 1,..., F T denote the corresponding marginal distributions on RD. Given F and t F t, define the second moment matrices = E[ >] 2 RTD TD and t = E[ t >t ] 2 RD D, and let r = rank() and rt = rank( t). Let M 2 Rr TD be a matrix whose rows form a basis of supp(F), and similarly let Nt 2 Rrt D be a matrix whose rows form a basis of supp(F t). Let N = diag(N1,..., NT), and define the Writing V =( V1| | VT), where Vt 2 Rrt d has rank dr, we can then construct the singular value decompositions Vt = Ut tW>t, with Ut 2 O(rt dt), t 2 Rdt dt and Wt 2 O(d dt), where dt = rank(Vt).


Spectral embedding for dynamic networks with stability guarantees

Neural Information Processing Systems

We consider the problem of embedding a dynamic network, to obtain time-evolving vector representations of each node, which can then be used to describe changes in behaviour of individual nodes, communities, or the entire graph. Given this open-ended remit, we argue that two types of stability in the spatio-temporal positioning of nodes are desirable: to assign the same position, up to noise, to nodes behaving similarly at a given time (cross-sectional stability) and a constant position, up to noise, to a single node behaving similarly across different times (longitudinal stability). Similarity in behaviour is defined formally using notions of exchangeability under a dynamic latent position network model. By showing how this model can be recast as a multilayer random dot product graph, we demonstrate that unfolded adjacency spectral embedding satisfies both stability conditions. We also show how two alternative methods, omnibus and independent spectral embedding, alternately lack one or the other form of stability.


Supplementary information for Learning Gaussian Mixtures with Generalised Linear Models Precise Asymptotics in High dimensions

Neural Information Processing Systems

This appendix presents the proof of the main technical result, Theorem 1. Throughout the whole proof, we assume that the set of conditions from Sec. 2 is verified. A.1 Required background In this Section, we give an overview of the main concepts and tools on approximate message passing algorithms which will be required for the proof. We start with some definitions that commonly appear in the approximate message-passing literature, see e.g. The main regularity class of functions we will use is that of pseudo-Lipschitz functions, which roughly amounts to functions with polynomially bounded first derivatives. We include the required scaling w.r.t. the dimensions in the definition for convenience. Since K will be kept finite, it can be absorbed in any of the constants. For example, the function f: Rn R,x7 1nkxk22 is pseudo-Lipshitz of order 2. Moreau envelopes and Bregman proximal operators -- In our proof, we will also frequently use the notions of Moreau envelopes and proximal operators, see e.g.


Learning Gaussian Mixtures with Generalised Linear Models: Precise Asymptotics in High-dimensions

Neural Information Processing Systems

Generalised linear models for multi-class classification problems are one of the fundamental building blocks of modern machine learning tasks. In this manuscript, we characterise the learning of a mixture of KGaussians with generic means and covariances via empirical risk minimisation (ERM) with any convex loss and regularisation. In particular, we prove exact asymptotics characterising the ERM estimator in high-dimensions, extending several previous results about Gaussian mixture classification in the literature. We exemplify our result in two tasks of interest in statistical learning: a) classification for a mixture with sparse means, where we study the efficiency of `1 penalty with respect to `2; b) max-margin multiclass classification, where we characterise the phase transition on the existence of the multi-class logistic maximum likelihood estimator for K >2. Finally, we discuss how our theory can be applied beyond the scope of synthetic data, showing that in different cases Gaussian mixtures capture closely the learning curve of classification tasks in real data sets.


What Knowledge Gets Distilled in Knowledge Distillation? Utkarsh Ojha Yuheng Li Anirudh Sundara Rajan Yingyu Liang Yong Jae Lee University of Wisconsin-Madison

Neural Information Processing Systems

Knowledge distillation aims to transfer useful information from a teacher network to a student network, with the primary goal of improving the student's performance for the task at hand. Over the years, there has a been a deluge of novel techniques and use cases of knowledge distillation. Yet, despite the various improvements, there seems to be a glaring gap in the community's fundamental understanding of the process. Specifically, what is the knowledge that gets distilled in knowledge distillation? In other words, in what ways does the student become similar to the teacher?


Supplementary material to Generalization Error Rates in Kernel Ridge Regression The Crossover from the Noiseless to Noisy Regime of the decays

Neural Information Processing Systems

A.1 Equations for Gaussian design In this Appendix we discuss the derivation of eqs. Exact asymptotic formulas for the excess prediction error of least-squares and ridge regression are a classic result in high-dimensional statistics, and have been derived in many different works [23, 32, 52, 53]. In this manuscript, we follow the presentation given in [25], which is particularly adapted to our derivation and has the advantage to hold rigorously at large but finite number of samples nand features p. We start by reviewing the formulas in [25]. Note that the risk considered in eq.