Learned mask refiners improve segmentation accuracy, but it is hard to tell how much of the improvement comes from explicit structure rather than from extra capacity, and whether it holds up when the mask generator or its error distribution changes. UnfoldCRF treats refinement as inference in a conditional random field over pixel labels and latent region variables. Its energy has a corrected unary term, learned local pairwise interactions, and image-conditioned latent-region consistency, with a null state that lets a region with weak label agreement withdraw from the consistency term; inference unrolls damped mean-field updates on this one energy. To isolate the effect of structure, we compare against recurrent black-box refiners that read the same inputs and receive the same parameter budget, stage count, and supervision. On COD10K, UnfoldCRF beats the strongest matched control by 1.0 Fβω point, improves all four COD metrics, and lowers the fraction of images made worse from 11.7% to 8.5%. Under a train-once protocol over five datasets and several mask sources, the 2.6M-parameter variant gains 4.2 mean ΔIoU against 2.0 for its control, and a variant built on frozen DINOv2 features matches the strongest foundation-model refiner with about a seventh of its resident parameters while staying ahead of its own control. On mask generators never seen in training, the gain is 2.0 Fβω points against 0.6 for the control. Zeroing individual messages shows where the corrections come from: the pairwise messages mostly fix boundaries, the region messages mostly fix non-boundary errors. Code and supporting materials will be publicly released.
Figures & tables
Figure 1: Refinement formulations. (a) Predefined CRFs use prescribed affinities and optional fixed regions. (b) Black-box refiners learn iterative mask updates without an explicit energy. (c) Reconstruction-derived unfolding alternates reconstruction and mask updates. (d) UnfoldCRF performs damped mean-field inference on a shared, image-conditioned segmentation energy with corrected unary, local pairwise, and latent-region terms.
Figure 2: Framework of UnfoldCRF. Top: I and Pinit define a conditional segmentation energy that is shared by all T stages and initializes Q0 ; only the damping coefficient αt varies by stage. Bottom left: the encoder constructs the corrected unary potential, symmetric local interactions, and multi-scale soft pixel–region incidences. Bottom right: stage t forms the region posterior Rt from Qt , computes the pairwise and mass-normalized latent-region messages, and combines them with the fixed unary term to obtain Qt+1 ; natural-parameter damping then produces Qt+1 .
Method
Resident params
BIG
VOC
DAVIS585
ECSSD
MSRA-B
Mean
Unrefined (IoU / BIoU)
n/a
78.3/70.1
66.7/60.1
80.1/83.0
81.4/70.2
75.2/61.9
76.3/69.1
CascadePSP ( Cheng et al., 2020 )
68M
+5.0/+6.3
+1.7/+0.7
− 1.3/ − 1.5
+0.6/+1.0
+0.9/+2.7
+1.4/+1.9
SegRefiner ( Wang et al., 2023 )
119M
+9.6 / +12.5
− 3.9/ − 3.1
− 10.9/ − 9.1
− 15.0/ − 21.4
− 10.7/ − 16.2
− 6.2/ − 7.4
DualSight ( Price et al., 2025 )
641M
+3.9/+4.6
+3.3/+6.3
+1.5/ − 0.6
+2.0/+4.5
+2.1/+6.8
+2.6/+4.3
SAMRefiner ( Lin et al., 2025 )
641M
+6.8/+9.5
+7.1/+9.7
+3.3/+2.0
+5.1/+9.7
+4.7/+10.4
+5.4/+8.3
PromptMoE ( Price et al., 2026 )
≥ 641M
+8.5/+11.0
+7.9/+10.4
+3.6 / +2.4
+6.0 /+10.7
+5.1/ +10.5
+6.2/+9.0
Table 1: General open refinement on the five-benchmark suite of Price et al. (2026) : Δ IoU / Δ BIoU (%) over the unrefined masks (absolute scores in the first row). Published baselines use the reported high-resolution variants under the same coarse-mask protocol. Bold: best per column. The last four rows are our runs, reported as means over 5 training seeds.
CHAMELEON
CAMO
COD10K
NC4K
Method
Source
M / Fβω / Eϕ / Sα
M / Fβω / Eϕ / Sα
M / Fβω / Eϕ / Sα
M / Fβω / Eϕ / Sα
SINet V2
TPAMI’21
.029/.792/.922/.890
.071/.733/.875/.822
.036/.668/.867/.820
.048/.769/.898/.848
+ SegRefiner
NeurIPS’23
.073/.619/.812/.777
.116/.583/.763/.736
.079/.487/.764/.714
.075/.635/.830/.781
+ Phoenix
ECCV’26
.030/.851/.915/.866
.079/.768/.834/.781
.034/ .766 /.872/.818
.049/ .822 /.883/.836
+ UMBD
arXiv’25
.027/ .853 / .955 /.892
.067/.774/.885/.829
.032/.733/.909/.825
.044/.807/.911/.849
+ UnfoldCRF-S
Ours
.026 /.849/.953/ .898
.066 / .775 / .890 / .834
.031 /.740/ .910 / .832
.043 /.812/ .915 / .855
Table 2: Concealed target refinement on COD under the UMBD ( Shen et al., 2025 ) protocol: M↓ / Fβω↑ / Eϕ↑ / Sα↑ on four test sets with SINetV2 ( Fan et al., 2021 ) and FEDER ( He et al., 2023a ) upstreams. SegRefiner ( Wang et al., 2023 ) and UMBD results are quoted from UMBD; Phoenix ( Kim and Hwang, 2026 ) is evaluated zero-shot on the same coarse masks, and our predictions use the same metric implementation. UMBD additionally uses upstream features.
Figure 3: Qualitative refinement. (a) General open refinement on representative examples. (b) Concealed target refinement on representative COD examples. Columns show four examples; rows show the image, initial mask, SegRefiner, SAMRefiner/UMBD, UnfoldCRF, and ground truth.
Operator
Params
M↓
Fβω↑
Eϕ↑
Sα↑
Harm ↓
Components
Upstream (FEDER)
n/a
.032
.713
.899
.822
n/a
Unary only
2.4M
.031
.725 ± .003
.903
.826
15.6 ± 0.6
Unary + learned pairwise
2.5M
.030
.736 ± .002
.906
.829
11.2 ± 0.5
Unary + learned latent regions
2.5M
.030
.738 ± .003
.907
.830
10.7 ± 0.5
Full UnfoldCRF-S ( T=5 )
2.6M
.029
.745 ± .002
.910
.834
8.5 ± 0.4
Table 3: Controlled comparison on COD10K with FEDER upstream (5 seeds; SD for Fβω and Harm). Component ablations share the encoder and training setting; structure controls use predefined pairwise or region structure; recurrent controls match the stage count and parameter budget. Harm is the percentage of images whose IoU decreases after refinement.
Figure 4: Efficiency, generalization, and mechanism. (a) Mean Δ IoU versus resident parameters; arrows indicate gains over matched recurrent controls at both scales. (b) Fβω gains on held-out generators. (c) Correction rates by error type under message interventions.
Figure 5: Stage-wise inference dynamics of one default checkpoint ( T=5 ) on COD10K: (a) Fβω of intermediate marginals with the initial-mask level as reference, (b) variational free energy relative to Q0 (test-set mean of the per-image difference), and (c) undamped fixed-point residual. Vertical dashed line: trained depth; shading: test-time extrapolation beyond T with αt:=αT−1 . Residuals are recorded before each update for t=0,…,2T−1 .
Figure 6: Representative failure cases with FEDER upstream. Columns show the image, initial mask, SegRefiner, UMBD, Phoenix, UnfoldCRF-S, and ground truth. Rows show a confidently mislabeled region, a missing object part, and a thin structure that remains difficult to recover.
official checkpoints (torchvision, Hugging Face); DAVIS585 masks from the dataset
VOC 2012 train
published where input provenance is matched; otherwise rerun
COD
SINetV2, FEDER
predictions regenerated from the official upstream checkpoints following the UMBD protocol
CAMO + COD10K train
published ( Shen et al., 2025 ) ; Phoenix rerun zero-shot
Controlled (COD10K, FEDER)
FEDER
as for COD
CAMO + COD10K train
our runs
Appendix
Table S1: Provenance of inputs and baseline numbers.
Variant
CHAMELEON
CAMO
NC4K
Unary only
.029/.833/.945/.891
.071/.748/.869/.804
.044/.803/.911/.850
Unary + learned pairwise
.028/.842/.948/.892
.069/.758/.872/.807
.043/.810/.913/.851
Unary + learned latent regions
.028/.840/.951/.894
.069/.755/.877/.811
.043/.808/.915/.853
Full UnfoldCRF-S
.026/.848/.953/.896
.067/.763/.879/.813
.042/.815/.917/.855
Without null state
.027/.841/.949/.892
.069/.754/.874/.809
.043/.809/.914/.852
Unary + fixed pairwise
.029/.836/.946/.891
.070/.752/.870/.806
.044/.806/.912/.851
Appendix
Table S2: Controlled variants on the other COD test sets ( M / Fβω / Eϕ / Sα ; FEDER upstream; mean over 5 seeds). Results on COD10K for the same variants are reported in Table 3 .
Despite significant advances in image segmentation, even state-of-the-art models produce masks with imperfect boundaries, semantic inconsistencies, and structural errors. Mask refinement addresses these limitations, yet current approaches rely on simplistic synthetic noise that fails to capture the complex error patterns of real segmentation models. We introduce Phoenix, a novel framework that leverages adversarial learning to generate semantically meaningful noise patterns and contrastive learning to model refinement relationships. Our approach consists of two key innovations: (1) Adversarial Mask Perturbation, which employs embedding attacks to create semantic-aware noise that mimics real segmentation errors, and (2) Contrastive Mask Refinement Learning, which establishes a tri-directional framework that ensures feature consistency within semantic regions while maintaining separation between classes. Experiments demonstrate that Phoenix significantly outperforms existing methods across diverse tasks, while consistently enhancing state-of-the-art segmentation models with substantial improvements. Our code and project page are publicly available at https://phoenix-eccv26.github.io.
Object-centric models often produce fragmented masks, boundary leakage, and incorrect region merging. We introduce Similarity-Shift Refinement (SSR), a training-free post-hoc method for improving object-centric masks with a frozen self-supervised Vision Transformer. SSR measures changes in pairwise patch similarity before and after self-attention value aggregation, retains positively strengthened relations, and constructs a sparse affinity graph. This graph propagates the initial soft slot assignments in a single refinement step, without retraining or modifying either model. Across natural-image, synthetic-video, and real-world-video benchmarks, SSR improves all-pixel Adjusted Rand Index in all 24 evaluated model-dataset combinations, with an average gain of 8.5 percentage points. Ablations show that value-space similarity shifts outperform query- and key-space variants as well as static Transformer affinities. However, texture-dense scenes may cause visually similar regions to be over-grouped. Overall, SSR provides a simple and transferable signal for training-free object-centric mask refinement.
Masked diffusion models (MDMs) generate discrete sequences by iterative denoising under an absorbing masking process. In standard masked diffusion, if a token remains masked after a reverse update, the model discards its clean-state prediction for that position. Thus, still-masked positions must be repeatedly inferred from the mask token alone. This design choice limits cross-step refinement. To address this limitation, this paper proposes a simple, yet effective, post-training adaptation for MDMs that conditions each denoising step on the model's own previous clean-state predictions. The resulting method, called Self-Conditioned Masked Diffusion Models (SCMDM), requires minimal architectural change, does not introduce a recurrent latent-state pathway, does not rely on an auxiliary reference model, and adds no extra denoiser evaluations during sampling. This is an important departure from partial self-conditioning approaches which requires expensive model training from scratch. In particular, the paper shows that partial self-conditioning, including the commonly used 50% dropout strategy for training self-conditioned models from scratch, is suboptimal in the post-training regime. Instead, once the model's self-generated clean-state estimates become informative, the specialization to refinement is preferable to mixing conditional and unconditional objectives. SCMDM is evaluated across multiple domains, demonstrating consistent improvement over vanilla MDM baselines, achieving nearly a 50% reduction in generative perplexity on OWT-trained models (42.89 to 23.72), alongside strong improvements in discretized image synthesis quality, small molecular generation, and enhanced fidelity in genomic distribution modeling.