Statistical Learning
between the correctness of autodiff systems and that of applications (e.g., gradient descent) built upon autodiff systems?
We thank the reviewers for their constructive and inspiring feedback. As we cannot see R2 (i.e., Reviewer #2), we respond to the reviews by R1, R3, and R4 only. The correctness of autodiff systems defined in the paper could be misleading to practitioners. We agree with the reviewers' points that (i) the correctness of the applications built upon autodiff systems is as important Also, we do not claim that our correctness condition is "the" Rather we are just suggesting "a" correctness condition that can serve as a reasonable (possibly minimal) We will clarify this limitation in the revised version of the paper. Here are detailed responses to the point (ii) on the applications mentioned in the reviews.
Export Reviews, Discussions, Author Feedback and Meta-Reviews
First provide a summary of the paper, and then address the following criteria: Quality, clarity, originality and significance. This paper proposes the Latent Case Model (LCM), a Bayesian approach to clustering in which clusters are represented by a prototype (a specific sample from the data) and feature subspaces (a binary subset of the variables signifying those features that are relevant to the class). The approach is presented as being a Bayesian, trainable version of the Case-Based Reasoning approach popular in AI, and is motivated by the ways such models have proved highly effective in explaining human decision making. The generative model (Figure 1) represents each item as coming from a mixture of S clusters, where each cluster is represented by a prototype and subspace (as above) and a function \phi which generates features matching those of the prototype with high probability for features in the subspace, and uniform features outside it. The model is thus similar in functionality to LDA but quite different in terms of its representation.
Export Reviews, Discussions, Author Feedback and Meta-Reviews
First provide a summary of the paper, and then address the following criteria: Quality, clarity, originality and significance. This paper proposes closed-form estimators for sparsity-structured graphical models, expressed as exponential family distributions, under high-dimensional settings. The paper is strong in both the methodology and theory aspect. The paper is well structured and written. The paper would be even better if the authors can discuss the connection between the method in the paper with previous methods which can provide close-form estimators for special MRFs such as tree-width=1 or planar MRF.