AITopics | rtformer

RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer

Neural Information Processing SystemsMar-18-2026, 13:22:54 GMT

Recently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in this field, due to the time-consuming computation mechanism of transformer. We propose RTFormer, an efficient dual-resolution transformer for real-time semantic segmenation, which achieves better trade-off between performance and efficiency than CNN-based models. To achieve high inference efficiency on GPU-like devices, our RTFormer leverages GPU-Friendly Attention with linear complexity and discards the multi-head mechanism. Besides, we find that cross-resolution attention is more efficient to gather global context information for high-resolution branch by spreading the high level knowledge learned from low-resolution branch. Extensive experiments on mainstream benchmarks demonstrate the effectiveness of our proposed RTFormer, it achieves state-of-the-art on Cityscapes, CamVid and COCOStuff, and shows promising results on ADE20K.

artificial intelligence, machine learning, proceedings, (7 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Machine Learning (0.41)

Add feedback

training

Neural Information Processing SystemsFeb-8-2026, 04:54:16 GMT

RTFormer is consist of several convolution blocks and RTFormerblocks,andRTFormerblockcontains differenttypes of attention. Table 2 shows the performance of RTFormer on ImageNet classification. The first three results of multi-head external attention are with r = [0.125,0.25,1]respectively. As illustrated in Table 3, we can find that multi-head self-attention achieves32.7 mIoU, which performs better than multi-head external attentions with different settings ofr. Multi-head external attention can achieve a good inference speed, which is benefit from its linear complexity and the design of sharing external parameter for multiple heads. However,theperformance ofmulti-headexternal attention is suboptimal, as the network capacity is limited by those designs.

artificial intelligence, external attention, multi-head external attention, (12 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Vision (0.30)

Add feedback

TimeSemantic

Neural Information Processing SystemsFeb-8-2026, 04:54:13 GMT

Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in this field, due to the time-consuming computation mechanism of transformer. We propose RTFormer, an efficient dual-resolution transformer for real-time semantic segmenation, which achievesbetter trade-off between performance and efficiency than CNN-based models.

artificial intelligence, machine learning, segmentation, (18 more...)

Neural Information Processing Systems

Genre: Research Report (0.46)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.94)

Add feedback

30e10e671c5e43edb67eb257abb6c3ea-Paper-Conference.pdf

Neural Information Processing SystemsAug-14-2025, 04:13:53 GMT

proceedings, segmentation, semantic segmentation, (12 more...)

Neural Information Processing Systems

Genre: Research Report (0.68)

Technology:

Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Natural Language (0.94)
Information Technology > Artificial Intelligence > Representation & Reasoning (0.68)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.46)

Add feedback

RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer

Neural Information Processing SystemsOct-10-2024, 13:20:49 GMT

Recently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in this field, due to the time-consuming computation mechanism of transformer. We propose RTFormer, an efficient dual-resolution transformer for real-time semantic segmenation, which achieves better trade-off between performance and efficiency than CNN-based models. To achieve high inference efficiency on GPU-like devices, our RTFormer leverages GPU-Friendly Attention with linear complexity and discards the multi-head mechanism. Besides, we find that cross-resolution attention is more efficient to gather global context information for high-resolution branch by spreading the high level knowledge learned from low-resolution branch. Extensive experiments on mainstream benchmarks demonstrate the effectiveness of our proposed RTFormer, it achieves state-of-the-art on Cityscapes, CamVid and COCOStuff, and shows promising results on ADE20K.

efficient design, real-time semantic segmentation, rtformer, (3 more...)

Neural Information Processing Systems

Technology: