Goto

Collaborating Authors

 Statistical Learning


Reviews: The Impact of Regularization on High-dimensional Logistic Regression

Neural Information Processing Systems

The authors study the limiting distribution of certain functionals of the penalized maximum likelihood estimator in regression. The paper contains nontrivial new extensions of the work of Sur and Candes in the unpenalized case, and is well-written and interesting. The reviews were mostly positive and the paper is in good shape.


Review for NeurIPS paper: Softmax Deep Double Deterministic Policy Gradients

Neural Information Processing Systems

Additional Feedback: I have the following questions for the authors to clarify and respond. For the bias definition in Theorems 3 and 4, is E [T (s')] also dependent on \theta {true}? If yes, would this be a reasonable assumption? 2. The authors showed that the proposed estimator can simultaneously reduce over- and under-estimation bias. Such results, however, definitely depend on the choice of \beta. Could you elaborate more on how to choose this parameter?


Review for NeurIPS paper: Softmax Deep Double Deterministic Policy Gradients

Neural Information Processing Systems

The reviewers appreciate the simple idea brought up in the paper and the experiments designed to understand its effect and the theoretical justification. Some reviewers did express concerns regarding the significance of the theoretical results and the concerns remain after the rebuttal. Please try to incorporate these feedback in your final draft.


Reviews: Sparse Logistic Regression Learns All Discrete Pairwise Graphical Models

Neural Information Processing Systems

This paper gives a simple and elegant algorithm for solving the long-studied problem of graphical model estimation (at least, in the case of pairwise MRFs, which includes the classic Ising model). The method uses a form of constrained logistic regression, which in retrospect, feels like the "right" way to solve this problem. The algorithm simply runs this constrained logistic regression method to learn the outgoing edges attached to each node. The proof is elegant and modular: first, based on standard generalization bounds, a sufficient number of samples allows minimization of the logistic loss function. Second, this loss is related to another loss function (the sigmoid of the inner product of the parameter vector with a sample from the distribution).


Reviews: Sparse Logistic Regression Learns All Discrete Pairwise Graphical Models

Neural Information Processing Systems

A clever new approach for learning pairwise graphical models with a gain of O(k) wrt previous work (k is the alphabet size). Results seem correct (though a detailed mathematical check was not performed) and important for the problem.


Reviews: First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise

Neural Information Processing Systems

For Reviewer #1's concern about making theory, I tend to be open-minded since I can not find solid evidence that the paper is making theory only. For Reviewer #4's comment about the over-claim of the result the paper proved, my take is follows. First, for many problems, the true local minima enjoys the flat basin. A famous example I have is the following paper: McGoff, Kevin A., et al. "The Local Edge Machine: inference of dynamic models of gene regulation." Second, the authors have explained the motivation of using the Levy process to model the noise.



Reviews: Global Convergence of Least Squares EM for Demixing Two Log-Concave Densities

Neural Information Processing Systems

The task of demixing a balanced two log-concave densities with the same and known covariance matrix seems restrictive. The authors did not discuss whether the results apply to unbalanced mixtures of two log-concave densities, nor did they discuss the implication of requiring covariances for the two being the same or the knowledge of covariances. The works builds on previous works on global convergence guarantees via EM for a balanced 2 GMM (or mixture of 2 truncated Gaussians) with known covariance, as well as following up works on mixture of 2 linear regressions and mixtures of 2 1-d laplacian distributions. The methodological contribution is to modify the M-step. However, the technical difficulty of proving convergence or providing finite-sample analysis compared with previous works is unclear.


Reviews: Foundations of Comparison-Based Hierarchical Clustering

Neural Information Processing Systems

In this work the authors study hierarchical clustering under quadruplet comparison framework. The authors show that single and complete linkages are inherently comparison based and propose two variants of average linkage clustering exploiting quadruplet comparison. Exact hierarchy recovery guarantee is provided under planted hierarchical partition model and empirical evaluation is provided. The meaning of the variables \mu, \delta etc are hard to interpret from the description. They have been nicely summarized (and explained) in the appendix A.1.


Reviews: Foundations of Comparison-Based Hierarchical Clustering

Neural Information Processing Systems

The authors have proposed two variants of average linkage hierarchical clustering using quadruplet comparison framework. Theoretical results of hierarchy recovery is established under a suitable model. The reviewers are in agreement that the results are new and important. The authors should incorporate the suggestions made by the reviewers to further strengthen the paper.