Country
TransMIL: TransformerbasedCorrelatedMultiple InstanceLearningforWholeSlide ImageClassification
However, the current MIL methods are usually based on independent and identical distribution hypothesis, thus neglect the correlation among different instances. To address this problem, we proposed a new framework, called correlated MIL, and provided a proof for convergence. Based on this framework, we devised a Transformer based MIL (TransMIL), which explored both morphological and spatial information. The proposed TransMIL can effectively deal with unbalanced/balanced and binary/multiple classification with great visualization and interpretability.
ImprovedCoresetsforEuclideank-Means
In the most general setting, a coreset compresses the data set in such a way that for any set of previously specified candidate queries, the cost of evaluating the query and the cost of the coreset are similar,up to an arbitrarysmalldistortion. A popular subject in coreset literature is the Euclideank-means problem.
11fc8c98b46d4cbdfe8157267228f7d7-Supplemental-Conference.pdf
We follow most of the settings in Uni-Perceiver [93]: cross-entropy loss with label smoothing of 0.1 is adopted for all tasks, and the negative samples for retrieval tasks are only from the local batch in the current GPU. We also apply the same data augmentation techniques as Uni-Perceiver [93] to image and video modalities to avoid overfitting. There are some setting changes to improve the training stability of the original Uni-Perceiver. Following [102], a uniform drop rate for stochastic depth is used across all encoder layers and are adapted according to the model size. Additionally, LayerScale [101] is used to facilitate the convergence of Transformer training, and the same initialization of10 3 is set to all models for simplicity.
Appendix: RemodelSelf-AttentionwithGaussian KernelandNyströmMethod
Figure 1: Validation loss changes for50k steps. Consider a finite sequence{Xk} of independent, random, self-adjoint matrices with dimensionn. For a certainn-by-n orthogonal matrixH (HHT is a diagonal matrix) and ann-by-d uniform sub-sampling matrixS (as defined in Definition 1 in the main paper), we denote the sketching matrixΠ:= nS.WeaimtoshowHΠΠTHT cansatisfy(12,δ)-MApropertyforHHT bythe followinglemma. The first inequality of the preceding display holds due to the fact thatH is an orthogonal matrix. It is easy to check that C C =B(I PΠ)BT.
Locally-AdaptiveNonparametricOnlineLearning: SupplementaryMaterial
In case of generic convex losses, we use the more complex parameterless algorithm AdaNormalHedge. The following theorem states a slightly more general bound that holds for anyη-exp-concave loss function (for completeness,theproofisgiveninAppendixD). Nownotethatalthough the algorithm is actually initialized withw1,i = 1, Lemma 1 shows that the regret remains the same if we assume the algorithm is initialized withwE1. Suppose that Algorithm 5 is run using predictions and updates provided by AdaNormalHedge. Asinourlocally-adaptive setting node experts are local learners,byi,t should be viewed as the prediction of the local online learning algorithm sitting at nodeiof the tree.