Serpent: Scalable and Efficient Image Restoration via Multi-scale Structured State Space Models
Sepehri, Mohammad Shahab, Fabian, Zalan, Soltanolkotabi, Mahdi
–arXiv.org Artificial Intelligence
Image restoration is aimed at recovering a clean image from its degraded counterpart, encompassing crucial tasks such as superresolution [11, 22], deblurring [19, 27], inpainting [25, 14] and JPEG compression artifact removal [3]. End-to-end deep learning techniques that directly learn the mapping from corrupted images to their clean counterparts are the current state-of-the-art in most image recovery tasks. The careful design of such architectures has attracted considerable attention in recent years, and is crucial for the performance and efficiency of image restoration methods. Architectures composed of convolutional building blocks have achieved great success in a multitude of image restoration problems [15, 20] thanks to their compute efficiency. However, convolutional neural networks (CNNs) are limited in low-level vision tasks by two key weaknesses. First, convolutional filters are content-independent, that is different image regions are processed by the same filter. Second, convolutions have limited capability to model long-range dependencies due to the small size of kernels, requiring exceedingly deeper architectures to increase the receptive field. More recently, Transformer architectures such as the Vision Transformer [2], have shown enormous potential in a variety of vision problems, including dense prediction tasks such as image restoration [26, 23, 12, 28]. Vision Transformers split the image into non-overlapping patches, and process the patches in an embedded token representation.
arXiv.org Artificial Intelligence
May-29-2024
- Country:
- North America > United States
- California > Los Angeles County > Los Angeles (0.29)
- Africa > Central African Republic
- Ombella-M'Poko > Bimbo (0.04)
- North America > United States
- Genre:
- Research Report (1.00)
- Industry:
- Health & Medicine (0.48)
- Technology: