Reviews: Understanding Weight Normalized Deep Neural Networks with Rectified Linear Units

Neural Information Processing Systems 

This paper essentially seems to address 2 questions (a) Under certain natural weight constraints what is the Rademacher complexity upperbound for nets when we allow for bias vectors in each layer? There is a section 4.1 in the paper which is about generalization bounds for nets doing regression. Let me first say at the outset that the writing of the paper seems extremely bad and many of the crucial steps in the proofs look unfollowable. As it stands this paper is hardly fit to be made public and needs a thorough rewriting! If that is the entire point then why is this interesting?