Goto

Collaborating Authors

 Statistical Learning



Scalable Diverse Model Selection for Accessible Transfer Learning Supplemental Material (Appendix)

Neural Information Processing Systems

We display full results for all methods here. This means that the source feature quality doesn't matter nearly as much. Since source feature quality is the only metric these methods use to predict transfer performance, they do poorly here. In Tab. 1, we display the Pearson Correlation for each target dataset individually. We also include results for additional baselines and skews of existing methods.





A The Estimator null A X W)

Neural Information Processing Systems

A.2 Proof of Theorem 1 To prove Theorem 1, we assume that G Proof of Lemma 1. Let's first rewrite Equation (4) as null null By Lemma 1, linearity of expectation and knowing that each RWT is independent from the other tours by the Strong Markov Property, Theorem 1 holds. MHM-GNN can recover edge-based models where representations don't use graph-wide However, on Rent the Runway we see the raw features achieving the highest performance. That is, structural information does not seem to be relevant to this specific task. All hyperparameters were chosen to minimize training loss. For k = 5, we used a minibatch of size 5 in all datasets.





Supplementary Material Training for the Future: A Simple Gradient Interpolation Loss to Generalize Along Time

Neural Information Processing Systems

In the main text, many algorithmic details were omitted and only discussed briefly. A.1 Dataset Details We expand upon the seven datasets used for our experiments in this section. The task is multi-class classification with a heavy class imbalance. It has 8 features including price, day of the week and units transferred. We discard instances with missing values.