Technology
dbc4d84bfcfe2284ba11beffb853a8c4-AuthorFeedback.pdf
Note that the theoretical equivalence requires near-zero initialization, gradient flow (small5 learning rate),and alargenumber ofchannels. These require significant computation resource. One of the advantages of kernel methods is that they requirelittle21 computationon a small dataset, which is a very appealing feature for architecture search.
Diminishing Returns Shape Constraints for Interpretability and Regularization
Maya Gupta, Dara Bahri, Andrew Cotter, Kevin Canini
Similarly, a model that predicts the time it will take a customer to grocery shop should decrease in the number of cashiers, but each addedcashierreduces average wait time by less. In both cases, we would like to be able to incorporate this prior knowledge by constraining the machine learned model's output to have a diminishing returns response to the size of the apartment or number of cashiers.
db98dc0dbafde48e8f74c0de001d35e4-AuthorFeedback.pdf
Re Eq. (3): indeed, we defineθi only later (line 157); we'll fix this, thanks. As requested, we here add another evaluation using34 t-SNE visualization (Figure 1). We do not assume data is35 available at large scale (though wecan and do handle36 large data sets as well);e.g., UCR contains 85 different37 datasets, manyofwhich include onlyfewexemplars per-38 class('ECGFiveDays',Fig1.,mainpaper,hadonly 1039 samples per class). We believe our experiments section40 was extensive and thorough, and we refer the reviewer41 to our supmat which includes more analyses and results.42