cs.CVMay 25, 2026

SP-MoMamba: Superpixel-driven Mixture of State Space Experts for Efficient Image Super-Resolution

Authors: Wenbin ZouYawen CuiYi WangLap-Pui ChauLiang ChenJinshan PanHuiping ZhuangGuanbin Li

Abstract

State space models (SSMs) have emerged as a powerful paradigm for efficient single-image super-resolution (SR) due to their linear complexity and long-range modeling capabilities. However, existing Mamba-based methods typically rely on data-agnostic rigid scanning, which reshapes 2D images into 1D sequences over a fixed grid, inevitably disrupting spatial-semantic topology and introducing artifacts. Inspired by the \textbf{Gestalt perceptual grouping theory}, we propose \textbf{SP-MoMamba}, a superpixel-driven mixture of state space experts designed for content-aware SR. Our core idea is to transform the traditional rigid scanning into a \textbf{semantic-level interaction} by treating superpixels as fundamental units. Specifically, we introduce the \textbf{Superpixel-driven State Space Model (SP-SSM)}, which compresses semantically homogeneous regions into high-order tokens to preserve global topological consistency. To address the conflict between fixed scanning scales and diverse semantic granularities, we develop the \textbf{Multi-Scale Superpixel Mixture of State Space Experts (MSS-MoE)}. This module utilizes a dynamic routing mechanism to adaptively assign scale-specific experts, effectively capturing multi-scale textures while reducing computational redundancy. Furthermore, to prevent the loss of high-frequency details during global abstraction, we introduce a \textbf{Local Spatial Modulation Expert (LSME)} to complement the global modeling, ensuring a precise reconstruction of sharp edges and fine structures. Extensive experiments on standard benchmarks demonstrate that SP-MoMamba achieves superior reconstruction fidelity and a more favorable efficiency-performance trade-off compared to state-of-the-art efficient SR methods.

Explore similar work

Mar 25, 2025cs.CV

Keyframe-Centric State-Space Modeling for Burst Image Super-Resolution

Burst image super-resolution (BISR) reconstructs a high-resolution keyframe by aggregating complementary sub-pixel evidence from a short burst of low-resolution frames. Existing methods often process all burst frames with heavy backbones or maintain deep cross-frame interaction throughout the network, leading to redundant computation on non-key frames and limiting scalability with burst length. In this work we propose BurstMamba, a BISR architecture built on a simple principle: allocate most compute to reconstructing the keyframe, and process the burst primarily to extract sub-pixel priors. To this end, BurstMamba decouples BISR into a high-capacity keyframe super-resolution stream and a lightweight burst stream that interacts with it only through stage-wise residual injection. To improve burst-to-keyframe transfer, we introduce Gather -> Aggregate -> Scatter (GAS), which uses correspondence only for cross-frame message passing while preserving native-view features through a residual connection, and a wavelet-conditioned state update that biases selective routing toward high-frequency regions. Across SyntheticSR, RealBSR-RAW, and RealBSR-RGB, BurstMamba achieves state-of-the-art results. Extensive ablations show that the proposed compute separation, GAS, and wavelet-conditioned routing each contribute to the model's accuracy and scaling.
Ozan Unal, Steven Marty, Dengxin Dai
Jun 27, 2026cs.CV

FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution

Diffusion prior-based methods have shown impressive results in real-world image super-resolution (ISR), yet two key challenges persist: balancing pixel-level fidelity with semantic quality, and adapting to diverse degradations. Existing dual-branch approaches freeze the pixel module during semantic training, but the semantic branch can still expand capacity within the pixel subspace, precluding genuine perceptual improvement. Moreover, using a single static adapter cannot generalize across heterogeneous real-world corruptions. To address both issues, we propose FreqOrtho-SR, which comprises: Freq\textbf{Freq}uency-guided Mixture of LoRA Experts (FreqMoE), it routes inputs to specialized experts via a non-parametric FFT-based degradation-feature extractor that encodes frequency-domain signatures, enabling stable and interpretable specialization across corruption types; and Ortho\textbf{Ortho}gonal Gradient Projection (OGP), which reframes the dual-objective optimization as a subspace-constrained problem: by extracting the pixel-fidelity subspace via SVD on combined expert weight deltas and projecting semantic gradients onto its null space, OGP guarantees orthogonality between the two objectives, enabling genuinely complementary learning without mutual interference. Experiments show that FreqOrtho-SR achieves competitive overall performance and a strong fidelity-perception trade-off across multiple benchmarks with efficient single-step inference. The source code of our method can be found at \href\href{https://github.com/sonhm3029/FreqOrtho-SR}{\texttt{sonhm3029/FreqOrtho-SR}}.
Minh Son Hoang, Dinh Phu Tran, Quyen Nguyen Duc +2
May 17, 2026cs.CV

EchoSR: Efficient Context Harnessing for Lightweight Image Super-Resolution

Image super-resolution (SR) aims to reconstruct high-quality, high-resolution (HR) images from low-resolution (LR) inputs and plays a critical role in various downstream applications. Despite recent advancements, balancing reconstruction fidelity and computational efficiency remains a fundamental challenge, particularly in resource-constrained scenarios. While existing lightweight methods attempt to expand receptive fields, many of them either incur substantial computational overhead, naively scale up kernel sizes, or lack mechanisms for coherent multi-scale integration, limiting their overall effectiveness and scalability. To address these limitations, we propose EchoSR, an efficient context-harnessing framework for lightweight image super-resolution, which unifies multi-scale receptive field modeling and hierarchical context fusion. EchoSR decouples feature learning into disentangled local, multi-scale, and global modeling stages through an efficient context-harnessing strategy, and further promotes seamless cross-scale integration via a cross-scale overlapping fusion mechanism. Extensive experiments have shown that EchoSR consistently outperforms state-of-the-art lightweight super-resolution methods across multiple benchmarks, while also achieving a faster speed (2×)(\sim 2\times). The source code is available at https://github.com/funnyWang-Echoes/EchoSR.
Hanli Zhao, Binhao Wang, Shihao Zhao +3