Feature pyramid methods, from FPN to BiFPN, have achieved strong performance in face detection by fusing multi-scale features. However, detecting faces under unconstrained conditions, such as small scale, occlusion, and extreme pose, remains difficult, as it requires global cross-scale dependencies that local fusion cannot model. State space models such as Mamba provide global context with linear complexity by scanning features as a sequence, and therefore offer a promising direction for this problem. Nevertheless, such a scan needs the two pyramid scales combined into a single feature map, and the way they are combined determines whether cross-scale structure is preserved. Summation collapses the two scales before the scan, so the scan has no cross-scale structure to exploit, while concatenation keeps both scales but at far higher cost. To address this, we propose \textbf{Weave Mamba Fusion (WMF)}, which interleaves two adjacent pyramid scales column by column so that each step of a horizontal bidirectional SS2D scan moves from one scale to the other. With partial-channel processing and parameter-free de-weaving, WMF enables efficient cross-scale interaction while preserving feature structure. Integrating WMF into every fusion node yields \textbf{WeaveBiFPN}, the neck of our \textbf{WeaveFace} detector. On WIDER FACE, WeaveFace achieves 91.41% mean AP with only 0.34M parameters and 1.16 GFLOPs, outperforming prior detectors under 0.5M parameters. Its largest gains are on the Hard subset, where it reaches 87.14% AP. The code is publicly available at https://github.com/dohun-mat/WeaveMambaFusion.
Figures & tables
Figure 1 : Feature map visualization of different feature pyramid networks. Given the input image, we visualize the fused features from (a) FPN, (b) PANet, (c) BiFPN, and (d) our WeaveBiFPN. Our WeaveBiFPN separates faces from the background more clearly than the others.
Figure 2 : Comparison of feature pyramid architectures. (a) FPN. (b) PANet. (c) BiFPN. (d) Our WeaveBiFPN. The dashed boxes enlarge the feature fusion in each design. Our WeaveBiFPN keeps the bidirectional structure of BiFPN but replaces its weighted fusion with Weave Mamba Fusion.
Figure 3 : Overview of WeaveFace .
Model
Backbone
Neck
Fusion
Easy
Medium
Hard
Avg. (%)
#Params (M)
GFLOPs
DSFD ‡ [ 13 ]
ResNet-152
FEM
Conv.
96.60
95.70
90.40
94.23
120.06
259.55
TinaFace ‡ [ 37 ]
ResNet-50
FPN
Conv.
97.00
96.30
93.40
95.57
37.98
172.95
TransEnc-R50 †
ResNet-50
FPN
Self-attention
93.03
92.89
88.56
91.49
33.72
32.92
RetinaFace ‡ [ 6 ]
ResNet-50
FPN
Conv.
96.70
96.10
91.40
94.73
29.50
37.59
SCRFD-10GF [ 9 ]
Basic Res
PANet
Conv.
95.93
94.95
90.81
93.90
3.86
9.98
FaceBoxes ‡ [ 34 ]
-
-
-
85.90
81.60
55.70
74.40
1.01
0.28
Table 1 : Comparison with state-of-the-art face detectors on the WIDER FACE validation set (multi-scale AP). Fusion denotes the operator family used for multi-scale feature fusion and context aggregation. Results are from original publications or the multi-scale evaluations in [ 9 ] ( ‡ ) and [ 11 ] ( § ). † : self-attention baseline (ResNet-50 backbone with a DETR Transformer encoder [ 3 ] ).
Method
AFW AP (%)
PASCAL AP (%)
#Params
MogFace [ 17 ]
99.85
99.32
85.26M
EfficientSRFace-L [ 24 ]
99.94
98.84
18.84M
SCRFD-0.5GF [ 9 ]
98.60
98.54
0.57M
SCRFD-1.0GF [ 9 ]
99.70
98.60
0.64M
SCRFD-2.5GF [ 9 ]
99.82
98.91
0.67M
FaceBoxes [ 34 ]
98.91
96.30
1.01M
Table 2 : Comparison on AFW and PASCAL Faces without fine-tuning, following the AP protocol of [ 19 ] . All models are trained solely on WIDER FACE. Bold : our model; underline : best result in each column.
Figure 4 : Qualitative comparison between EResFD (left) and WeaveFace (right) on the WIDER FACE Hard subset. Red boxes represent successful detections. Yellow boxes indicate faces missed by EResFD but correctly identified by WeaveFace, while green boxes highlight false positives generated by EResFD that are effectively mitigated by WeaveFace.
Fusion
Easy
Medium
Hard
Avg.
#Params (M)
FLOPs (G)
Lat. (ms)
Sum (w/ sep. conv.)
92.88
90.87
84.76
89.50
0.217
0.564
15.1
Cross-attention
93.47
91.55
86.46
90.49
0.577
11.228
23.7
WMF
94.33
92.76
87.14
91.41
0.344
1.159
23.3
Table 3 : Comparison of different fusion operators on the BiFPN topology. We evaluate our proposed WMF against standard summation and cross-attention baselines. Bold : our model; underline : best result in each column; Lat. : latency.
Topology
Fusion
Easy
Medium
Hard
Avg.
#Params (M)
FLOPs (G)
FPN
Sum
91.17
89.31
83.81
88.10
0.229
0.672
WMF
91.47
89.72
84.12
88.44
0.187
0.639
PANet
Concat
92.33
90.35
84.88
89.19
0.562
1.071
WMF
92.75
91.08
85.03
89.62
0.329
0.751
BiFPN
Sum
92.88
90.87
84.76
89.50
0.217
0.564
WMF
94.33
92.76
87.14
91.41
0.344
1.159
Table 4 : Generality of WMF across pyramid topologies. Each topology is paired with its standard fusion method (summation for FPN/BiFPN, concatenation for PANet) and the proposed WMF. Bold : WMF (ours); underline : best result in each column within each topology.
Fusion Operator
N
Easy
Medium
Hard
Avg.
#Params (M)
FLOPs (G)
Sum + SS2D
1
93.27
91.54
85.40
90.07
0.409
1.152
2
91.71
89.82
83.85
88.46
0.245
0.711
4
90.18
88.13
81.20
86.50
0.193
0.566
Concat + SS2D
1
95.14
93.79
88.40
92.44
1.080
2.649
2
93.49
91.57
85.81
90.29
0.510
1.350
4
93.22
91.00
84.47
89.56
0.344
0.925
Table 5 : Accuracy–efficiency trade-offs across different scale-combination schemes (summation, concatenation, and weaving) as the partial-channel divisor N increases. N determines the extent of lightweighting in SS2D, where only 1/N of channels are processed; thus, a larger N reduces model parameters and FLOPs at the expense of accuracy. Bold : our final configuration; underline : best result in each column.
s
Easy
Medium
Hard
Avg.
#Params (M)
FLOPs (G)
1
93.87
92.07
86.51
90.82
0.323
0.991
2
94.33
92.76
87.14
91.41
0.344
1.159
4
94.01
92.52
87.06
91.20
0.386
1.496
Table 6 : Effect of the scan count s in WeaveBiFPN, where s denotes the number of scan directions utilized by SS2D. Bold : our final configuration ( s=2 ); underline : best result in each column.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Easy
Medium
Hard
Avg.
#Params (M)
FLOPs (G)
EResNet
94.33
92.76
87.14
91.41
0.344
1.159
MobileNetV1-0.25
95.06
93.64
88.52
92.41
0.566
1.231
ResNet-50
96.29
95.44
91.20
94.31
29.756
36.275
ResNet-152
96.33
95.44
91.15
94.31
64.392
81.871
Appendix
Table 7 : WeaveFace with different backbones on the WIDER FACE validation set. The WeaveBiFPN neck ( N=2 , s=2 ), CBAM, and CCPM head are fixed, and only the backbone is changed. The same fusion design generalizes from lightweight to heavy backbones. Bold : our default backbone.
Level
Stride
Feature map
Anchor sizes (px)
P2
8
80×80
16, 32
P3
16
40×40
64, 128
P4
32
20×20
256, 512
Appendix
Table 8 : Anchor settings of WeaveFace. Feature-map sizes are given for the 640×640 training resolution.
COCO val2017
PASCAL VOC
Variant
AP
AP 50
AP 75
Params (M)
FLOPs (G)
mAP
Params (M)
FLOPs (G)
EfficientDet (baseline)
34.5
52.9
36.6
3.88
2.57
76.90
3.84
2.35
EfficientDet + WMF
36.2
53.6
38.2
4.25
3.21
78.63
4.21
2.99
Appendix
Table 9 : Results on COCO val2017 and PASCAL VOC datasets. Only WMF is added to the BiFPN neck of EfficientDet; all other settings are unchanged.
Current face video forgery detectors use wide or dual-stream backbones. We show that a single, lightweight fusion of two handcrafted cues can achieve higher accuracy with a much smaller model. Based on the Xception baseline model (21.9 million parameters), we build two detectors: LFWS, which adds a 1x1 convolution to combine a low-frequency Wavelet-Denoised Feature (WDF) with a phase-spectrum channel derived from Spatial-Phase Shallow Learning (SPSL), and LFWL, which merges WDF with Local Binary Patterns (LBP) in the same way. This extra module adds only 292 parameters, keeping the total at 21.9 million, smaller than F3Net (22.5 million) and less than half the size of SRM (55.3 million). Even with this minimal overhead, the fused models increase the average area under the curve (AUC) from 74.8% to 78.6% on FaceForensics++ and from 70.5% to 74.9% on DFDC-Preview, gains of 3.8% and 4.4% over the Xception baseline. They also consistently outperform F3Net, SRM, and SPSL in eight public benchmarks, without extra data or test-time augmentation. These results show that carefully paired, handcrafted features, combined through the lightweight fusion block, can provide competitive robustness at a significantly lower cost than comparable frequency-based detectors. Our findings suggest a need to reevaluate scale-driven design choices in face video forgery detection.
Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different network parameters need to be trained independently for each $N$-point'' benchmark dataset, and (2) a model trained on an N-point'' dataset reliably outputs only the N landmarks. In our work, we first conceptualize Face Part-Anchored Landmark Positions (FPALPs), wherein each landmark is treated as a progression value between zero (start) and one (end) along a face part's contour. Every landmark can be expressed in the FPALP format, irrespective of its source dataset, hence unlocking the ability to unify all $N$-point'' datasets into a single dataset. Secondly, we represent each landmark with an FPALP-based query, refine it progressively with a cross-modality decoder, and predict its coordinates based on the final representation. Our approach, called Unified Dynamic FLD, embodies these two design choices and streamlines the landmark detection pipeline by enabling (1) a single model to learn on any number of N-point'' datasets, and (2) yield any number of specific landmark predictions by loading the designated landmark queries at runtime. Extensive experiments on multiple benchmark datasets show that our method delivers these benefits while remaining competitive with, and in several cases outperforming existing state-of-the-art methods.
Sebastian Regalado, Varshanth R. Rao, Ruowei Jiang +2
University of Toronto · 2ModiFace · †Work completed while employed at ModiFace +1
Feature fusion networks are fundamental components in modern object detectors, aggregating multi-scale features to detect objects of varying sizes. However, directly fusing features from different pyramid levels often introduces semantic inconsistency due to their heterogeneous representations. In this paper, we propose Feature Interaction NEtwork (FINE), a lightweight semantic alignment module that refines low-level features via high-level contextual guidance using cross-level attention prior to fusion. To bridge the structural gap and ensure computational efficiency, we introduce an Alignment-Aware Token Sampling that aligns corresponding spatial regions across scales, reducing the attention complexity by an order of magnitude. The resulting attention weights generate a spatial-channel modulation map that is upsampled and applied to the low-level features via residual element-wise modulation. This mechanism ensures that the network selectively enhances semantically relevant pixels while preserving the sub-pixel localization accuracy necessary for dense prediction tasks. FINE is generally applicable to various detectors and consistently improves detection accuracy without compromising efficiency.
Hyungseop Lee, Jiho Lee, Woochul Kang
Department of Embedded Systems Engineering, Incheon National University, Yeonsu-gu 22012, South Korea