Statistical Learning
Fast Transformers with Clustered Attention Supplementary Material
Figure 1: Flow-chart demonstrating the compuation for clustered attention. For more details refer to 1.1 or 3.2 in the main paper. Work done at Idiap 34th Conference on Neural Information Processing Systems (NeurIPS 2020), V ancouver, Canada. We then present the flow chart demonstrating the same. This is followed by taking the weighted average of the 3 correponding values.
Duality-Induced Regularizer for Tensor Factorization Based Knowledge Graph Completion Supplementary Material
Theorem 1. Suppose that ห X In DB models, the commonly used p is either 1 or 2. When p = 2, DURA takes the form as the one in Equation (8) in the main text. If p = 1, we cannot expand the squared score function of the associated DB models as in Equation (4). Therefore, we choose p = 2 . 2 Table 2: Hyperparameters found by grid search. Suppose that k is the number of triplets known to be true in the knowledge graph, n is the embedding dimension of entities. That is to say, the computational complexity of weighted DURA is the same as the weighted squared Frobenius norm regularizer.