Statistical Learning
A Comparison with Other General MLCO Frameworks
We would also like to discuss the limitations of the approaches including ours. As shown in Tab. 4, the PPO-Single that serves as a baseline in our paper is designed following As shown in Tab. 4, NerRewritter is most general because it can be viewed as a learning-based local It is also worth noting that there are some problems that are beyond our knowledge to tackle, e.g. the expression simplify problem, and it may requires experts with specific domain We have discussed the model details of PPO-BiHyb in Sec. 4, and in this section, we discuss the DAG. Considering the structure of DAG, we design two GCNs: the first GCN processes the original DAG, and the second GCN processes the DAG with all edges reversed. The predicted doubly-stochastic matrix by SK is processed by considering the partial matching matrix. Graph-level features are obtained via attention pooling, which are fed to the critic net.
R1: Comparison with inexact methods Aligning with prior exact papers [10, 18], we focus on comparisons with exact
We thank all five reviewers for their detailed and incisive feedback. We tested AustereMH [16], an inexact method, on robust linear regression in Section 5.1 with We added this to the Appendix. This does not affect the properties of TunaMH. Our theorem doesn't have this assumption; it suggests that for MHSubLhd with given user-specified The impact is 3-fold: it (1) provides an upper bound on performance for algorithms of Algorithm 1's TunaMH); (3) suggests directions for developing new algorithms. To be significantly faster than TunaMH, we either need more assumptions about the problem or new stateful algorithms.
main remarks regarding baseline, scalability, complexity and the full batch setting in the following paragraphs
We thank the reviewers for the valuable comments and suggestions made. The reviewers' main concern is the lack of RQVI procedure led to computational instability). GLM, BNN) and five datasets (Boston, Fires, Life Expect., Frisk and Metro) with learning rate analysis. We do not claim that this method is suitable for high dimensional posteriors. It is accurate that the method will not be viable without this property.