Goto

Collaborating Authors

 Technology



A Experimental Details

Neural Information Processing Systems

The number of communication rounds is 100. The total number of clients is 4. A two-layer CNN architecture is used as the backbone model here. The local training batch size is 32. An SGD optimizer with a weight decay rate 5e-4 and a learning rate 0.01 is used. There are 20 clients in total.




42c8938e4cf5777700700e642dc2a8cd-AuthorFeedback.pdf

Neural Information Processing Systems

It also makes no assumptions on the sparseness of the transitions. Our experiments reflect this as well,25 as the transition probabilities are drawn from a uniform distribution with no sparseness assumptions and would be26 more difficult tothan sparse cases.



cf78a15772ec1a6aee9bbee2d2b382c3-Supplemental-Conference.pdf

Neural Information Processing Systems

Our first step is to prove the parameterization (Eq. 3) provides local attention after the Note that the weight and bias terms in theaboveformulation (Eq. Assume the position-based function at each head is learned to perform'hard attention' on one of its surrounding positions,i.e., an extreme semi-dynamic attention. To demonstrate this phenomenon, we plot and compare the impacts ofฮฆc and ฮฆp6 on ฮฆa in the middle and right of Fig. S4 and visualize learned position-based attentionฮฆp of iRPE in Fig. S5. As seen from Tab. S17, there exist noticeable performance gaps between the models (b, f, g, h) (withoutฮฆp)and(a,d,e,i)(withฮฆp). Without adaptiveattention (model (c)),ฮฆp imposes stronger locality onevery layer.


PeripheralVisionTransformer

Neural Information Processing Systems

As seen in Figure 1, we have high-resolution processing near the center ofourgaze,i.e.,central and para-central regions, toidentify highly-detailed visual elements such as geometric shapes, and low-level details.