Goto

Collaborating Authors

 Technology



42c8938e4cf5777700700e642dc2a8cd-AuthorFeedback.pdf

Neural Information Processing Systems

It also makes no assumptions on the sparseness of the transitions. Our experiments reflect this as well,25 as the transition probabilities are drawn from a uniform distribution with no sparseness assumptions and would be26 more difficult tothan sparse cases.



cf78a15772ec1a6aee9bbee2d2b382c3-Supplemental-Conference.pdf

Neural Information Processing Systems

Our first step is to prove the parameterization (Eq. 3) provides local attention after the Note that the weight and bias terms in theaboveformulation (Eq. Assume the position-based function at each head is learned to perform'hard attention' on one of its surrounding positions,i.e., an extreme semi-dynamic attention. To demonstrate this phenomenon, we plot and compare the impacts ofΦc and Φp6 on Φa in the middle and right of Fig. S4 and visualize learned position-based attentionΦp of iRPE in Fig. S5. As seen from Tab. S17, there exist noticeable performance gaps between the models (b, f, g, h) (withoutΦp)and(a,d,e,i)(withΦp). Without adaptiveattention (model (c)),Φp imposes stronger locality onevery layer.


PeripheralVisionTransformer

Neural Information Processing Systems

As seen in Figure 1, we have high-resolution processing near the center ofourgaze,i.e.,central and para-central regions, toidentify highly-detailed visual elements such as geometric shapes, and low-level details.






Table3: HumanActivity

Neural Information Processing Systems

As reviewers noticed, detailed description of the experiments is provided in the3 supplement. Wealso6 switched to a finer discretization of 1 minute on Physionet dataset, instead of 6 minutes. Modeling state-interactions via26 an ODE allows better generalization outside the training interval compared to directly modeling a function of time.27 Using an ODE-RNN as a decoder is a possible extension.