Towards Understanding Learning Representations: To What Extent Do Different Neural Networks Learn the Same Representation

Wang, Liwei, Hu, Lunjia, Gu, Jiayuan, Hu, Zhiqiang, Wu, Yue, He, Kun, Hopcroft, John

Neural Information Processing Systems 

It is widely believed that learning good representations is one of the main reasons for the success of deep neural networks. Although highly intuitive, there is a lack of theory and systematic approach quantitatively characterizing what representations do deep neural networks learn. In this work, we move a tiny step towards a theory and better understanding of the representations. Specifically, we study a simpler problem: How similar are the representations learned by two networks with identical architecture but trained from different initializations. We develop a rigorous theory based on the neuron activation subspace match model.