Learned mask refiners improve segmentation accuracy, but it is hard to tell how much of the improvement comes from explicit structure rather than from extra capacity, and whether it holds up when the mask generator or its error distribution changes. UnfoldCRF treats refinement as inference in a conditional random field over pixel labels and latent region variables. Its energy has a corrected unary term, learned local pairwise interactions, and image-conditioned latent-region consistency, with a null state that lets a region with weak label agreement withdraw from the consistency term; inference unrolls damped mean-field updates on this one energy. To isolate the effect of structure, we compare against recurrent black-box refiners that read the same inputs and receive the same parameter budget, stage count, and supervision. On COD10K, UnfoldCRF beats the strongest matched control by 1.0 Fβω point, improves all four COD metrics, and lowers the fraction of images made worse from 11.7% to 8.5%. Under a train-once protocol over five datasets and several mask sources, the 2.6M-parameter variant gains 4.2 mean ΔIoU against 2.0 for its control, and a variant built on frozen DINOv2 features matches the strongest foundation-model refiner with about a seventh of its resident parameters while staying ahead of its own control. On mask generators never seen in training, the gain is 2.0 Fβω points against 0.6 for the control. Zeroing individual messages shows where the corrections come from: the pairwise messages mostly fix boundaries, the region messages mostly fix non-boundary errors. Code and supporting materials will be publicly released.
Figures & tables
Figure 1: Refinement formulations. (a) Predefined CRFs use prescribed affinities and optional fixed regions. (b) Black-box refiners learn iterative mask updates without an explicit energy. (c) Reconstruction-derived unfolding alternates reconstruction and mask updates. (d) UnfoldCRF performs damped mean-field inference on a shared, image-conditioned segmentation energy with corrected unary, local pairwise, and latent-region terms.
Figure 2: Framework of UnfoldCRF. Top: I and Pinit define a conditional segmentation energy that is shared by all T stages and initializes Q0 ; only the damping coefficient αt varies by stage. Bottom left: the encoder constructs the corrected unary potential, symmetric local interactions, and multi-scale soft pixel–region incidences. Bottom right: stage t forms the region posterior Rt from Qt , computes the pairwise and mass-normalized latent-region messages, and combines them with the fixed unary term to obtain Qt+1 ; natural-parameter damping then produces Qt+1 .
Method
Resident params
BIG
VOC
DAVIS585
ECSSD
MSRA-B
Mean
Unrefined (IoU / BIoU)
n/a
78.3/70.1
66.7/60.1
80.1/83.0
81.4/70.2
75.2/61.9
76.3/69.1
CascadePSP ( Cheng et al., 2020 )
68M
+5.0/+6.3
+1.7/+0.7
− 1.3/ − 1.5
+0.6/+1.0
+0.9/+2.7
+1.4/+1.9
SegRefiner ( Wang et al., 2023 )
119M
+9.6 / +12.5
− 3.9/ − 3.1
− 10.9/ − 9.1
− 15.0/ − 21.4
− 10.7/ − 16.2
− 6.2/ − 7.4
DualSight ( Price et al., 2025 )
641M
+3.9/+4.6
+3.3/+6.3
+1.5/ − 0.6
+2.0/+4.5
+2.1/+6.8
+2.6/+4.3
SAMRefiner ( Lin et al., 2025 )
641M
+6.8/+9.5
+7.1/+9.7
+3.3/+2.0
+5.1/+9.7
+4.7/+10.4
+5.4/+8.3
PromptMoE ( Price et al., 2026 )
≥ 641M
+8.5/+11.0
+7.9/+10.4
+3.6 / +2.4
+6.0 /+10.7
+5.1/ +10.5
+6.2/+9.0
Table 1: General open refinement on the five-benchmark suite of Price et al. (2026) : Δ IoU / Δ BIoU (%) over the unrefined masks (absolute scores in the first row). Published baselines use the reported high-resolution variants under the same coarse-mask protocol. Bold: best per column. The last four rows are our runs, reported as means over 5 training seeds.
CHAMELEON
CAMO
COD10K
NC4K
Method
Source
M / Fβω / Eϕ / Sα
M / Fβω / Eϕ / Sα
M / Fβω / Eϕ / Sα
M / Fβω / Eϕ / Sα
SINet V2
TPAMI’21
.029/.792/.922/.890
.071/.733/.875/.822
.036/.668/.867/.820
.048/.769/.898/.848
+ SegRefiner
NeurIPS’23
.073/.619/.812/.777
.116/.583/.763/.736
.079/.487/.764/.714
.075/.635/.830/.781
+ Phoenix
ECCV’26
.030/.851/.915/.866
.079/.768/.834/.781
.034/ .766 /.872/.818
.049/ .822 /.883/.836
+ UMBD
arXiv’25
.027/ .853 / .955 /.892
.067/.774/.885/.829
.032/.733/.909/.825
.044/.807/.911/.849
+ UnfoldCRF-S
Ours
.026 /.849/.953/ .898
.066 / .775 / .890 / .834
.031 /.740/ .910 / .832
.043 /.812/ .915 / .855
Table 2: Concealed target refinement on COD under the UMBD ( Shen et al., 2025 ) protocol: M↓ / Fβω↑ / Eϕ↑ / Sα↑ on four test sets with SINetV2 ( Fan et al., 2021 ) and FEDER ( He et al., 2023a ) upstreams. SegRefiner ( Wang et al., 2023 ) and UMBD results are quoted from UMBD; Phoenix ( Kim and Hwang, 2026 ) is evaluated zero-shot on the same coarse masks, and our predictions use the same metric implementation. UMBD additionally uses upstream features.
Figure 3: Qualitative refinement. (a) General open refinement on representative examples. (b) Concealed target refinement on representative COD examples. Columns show four examples; rows show the image, initial mask, SegRefiner, SAMRefiner/UMBD, UnfoldCRF, and ground truth.
Operator
Params
M↓
Fβω↑
Eϕ↑
Sα↑
Harm ↓
Components
Upstream (FEDER)
n/a
.032
.713
.899
.822
n/a
Unary only
2.4M
.031
.725 ± .003
.903
.826
15.6 ± 0.6
Unary + learned pairwise
2.5M
.030
.736 ± .002
.906
.829
11.2 ± 0.5
Unary + learned latent regions
2.5M
.030
.738 ± .003
.907
.830
10.7 ± 0.5
Full UnfoldCRF-S ( T=5 )
2.6M
.029
.745 ± .002
.910
.834
8.5 ± 0.4
Table 3: Controlled comparison on COD10K with FEDER upstream (5 seeds; SD for Fβω and Harm). Component ablations share the encoder and training setting; structure controls use predefined pairwise or region structure; recurrent controls match the stage count and parameter budget. Harm is the percentage of images whose IoU decreases after refinement.
Figure 4: Efficiency, generalization, and mechanism. (a) Mean Δ IoU versus resident parameters; arrows indicate gains over matched recurrent controls at both scales. (b) Fβω gains on held-out generators. (c) Correction rates by error type under message interventions.
Figure 5: Stage-wise inference dynamics of one default checkpoint ( T=5 ) on COD10K: (a) Fβω of intermediate marginals with the initial-mask level as reference, (b) variational free energy relative to Q0 (test-set mean of the per-image difference), and (c) undamped fixed-point residual. Vertical dashed line: trained depth; shading: test-time extrapolation beyond T with αt:=αT−1 . Residuals are recorded before each update for t=0,…,2T−1 .
Figure 6: Representative failure cases with FEDER upstream. Columns show the image, initial mask, SegRefiner, UMBD, Phoenix, UnfoldCRF-S, and ground truth. Rows show a confidently mislabeled region, a missing object part, and a thin structure that remains difficult to recover.
official checkpoints (torchvision, Hugging Face); DAVIS585 masks from the dataset
VOC 2012 train
published where input provenance is matched; otherwise rerun
COD
SINetV2, FEDER
predictions regenerated from the official upstream checkpoints following the UMBD protocol
CAMO + COD10K train
published ( Shen et al., 2025 ) ; Phoenix rerun zero-shot
Controlled (COD10K, FEDER)
FEDER
as for COD
CAMO + COD10K train
our runs
Appendix
Table S1: Provenance of inputs and baseline numbers.
Variant
CHAMELEON
CAMO
NC4K
Unary only
.029/.833/.945/.891
.071/.748/.869/.804
.044/.803/.911/.850
Unary + learned pairwise
.028/.842/.948/.892
.069/.758/.872/.807
.043/.810/.913/.851
Unary + learned latent regions
.028/.840/.951/.894
.069/.755/.877/.811
.043/.808/.915/.853
Full UnfoldCRF-S
.026/.848/.953/.896
.067/.763/.879/.813
.042/.815/.917/.855
Without null state
.027/.841/.949/.892
.069/.754/.874/.809
.043/.809/.914/.852
Unary + fixed pairwise
.029/.836/.946/.891
.070/.752/.870/.806
.044/.806/.912/.851
Appendix
Table S2: Controlled variants on the other COD test sets ( M / Fβω / Eϕ / Sα ; FEDER upstream; mean over 5 seeds). Results on COD10K for the same variants are reported in Table 3 .