Industry
cf78a15772ec1a6aee9bbee2d2b382c3-Supplemental-Conference.pdf
Our first step is to prove the parameterization (Eq. 3) provides local attention after the Note that the weight and bias terms in theaboveformulation (Eq. Assume the position-based function at each head is learned to perform'hard attention' on one of its surrounding positions,i.e., an extreme semi-dynamic attention. To demonstrate this phenomenon, we plot and compare the impacts ofฮฆc and ฮฆp6 on ฮฆa in the middle and right of Fig. S4 and visualize learned position-based attentionฮฆp of iRPE in Fig. S5. As seen from Tab. S17, there exist noticeable performance gaps between the models (b, f, g, h) (withoutฮฆp)and(a,d,e,i)(withฮฆp). Without adaptiveattention (model (c)),ฮฆp imposes stronger locality onevery layer.