Image Analysis and Stereology · Published 2025-02-05 · DOI 10.5566/ias.3399
Deep learning methods have demonstrated significant advancements in single image super-resolution (SISR), with Transformer-based models frequently outperforming CNN-based counterparts in performance. However, due to the self-attention mechanism in Transformers, achieving lightweight models remains challenging compared to CNN-based approaches. In this paper, we propose a lightweight Transformer model termed Multi-Branch Mixer Transformer (MBMT) for SR. The design of MBMT is motivated by two main considerations: while self-attention excels at capturing long-range dependencies in features, it struggles with extracting local features. Secondly, the quadratic complexity of self-attention forms a significant challenge in building lightweight models. To address these problems, we propose a Multi-Branch Token Mixer (MBTM) to extract richer global and local information. Compared to other Transformer-based SR networks, MBTM achieves a balance between capturing global information and reducing the computational complexity of self-attention through its compact multi-branch structure. Specifically, MBTM consists of three parts: shifted window attention, depthwise convolution, and active token mixer. This multi-branch structure handles both long-range dependencies and local features simultaneously, enabling us to achieve excellent SR performance with just a few stacked modules. Experimental results demonstrate that MBMT achieves competitive performance while maintaining model efficiency compared to SOTA methods.
Abstract from DOAJ. Public domain (CC0 1.0).
Read the article at the publisher →