From Routing Signals to Selective Review: Visual regrounding in MoE VLMs
Organizations: University of California, Los Angeles · Peking University
Abstract
Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, hidden states, or uncertainty measures. We present the first framework to leverage internal routing decisions in Mixture-of-Experts (MoE) VLMs to detect target absence before generation and guide selective correction. We extract target-token routing probabilities from Qwen3-VL-30B-A3B-Instruct and Gemma-4-26B-A4B-it, train a separate L2-regularized linear detector for each model, and use its predictions to selectively invoke a target-aware review prompt. Using routing alone, the Qwen and Gemma detectors achieve ROC-AUCs of 0.9988 and 0.9956 on GQA-Inpaint and retain 0.8095 and 0.7781 on the external OBER dataset, respectively. The resulting routing-gated policy improves end-to-end accuracy on GQA-Inpaint and OBER by +22.25% and +12.17% for Qwen, and by +13.42% and +1.39% for Gemma, without modifying model weights. Further analysis shows that the signal is localized to the target-object token, emerges in early MoE layers, and is distributed across partially substitutable experts. Although cross-dataset threshold shifts require recalibration, false-positive review causes limited harm overall, suggesting that intervention risk can be controlled through joint selection of the detector threshold and review prompt. Overall, we show that routing probabilities alone preserve actionable information about visual perception, allowing computation already produced by an MoE VLM to support low-cost detection and selective visual regrounding.
Figures & tables
| Representation | ROC-AUC | Accuracy |
|---|---|---|
| Input embedding | 0.5000 | 50.00% |
| Mean of hidden layers | 0.9979 | 98.06% |
| Last hidden layer (L47) | 0.9980 | 98.16% |
| All-layer routing | 0.9988 | 98.54% |
| Dataset | Target | Standard | Always | Selective |
|---|---|---|---|---|
| GQA | Absent | 37.17% | 83.17% | 82.50% |
| GQA | Present | 96.33% | 89.33% | 95.50% |
| GQA | Total | 66.75% | 86.25% | 89.00% |
| OBER | Absent | 66.78% | 95.83% | 95.83% |
| OBER | Present | 100.00% | 94.78% | 95.30% |
| OBER | Total | 83.39% | 95.30% | 95.57% |
| Rank | Expert | Mean weight |
|---|---|---|
| 1 | L04.E067 | |
| 2 | L02.E013 | |
| 3 | L45.E063 | |
| 4 | L02.E055 | |
| 5 | L13.E018 | |
| 6 | L02.E089 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Cohort | Pairs | Images | Role |
| GQA training, automatic QC | 20,000 | 40,000 | Probe fitting |
| GQA validation, automatic QC | 5,446 | 10,892 | Model and threshold selection |
| GQA held-out test, automatic QC | 5,776 | 11,552 | Frozen detector test |
| GQA manually audited review cohort | 120 | 240 | Same-domain policy evaluation |
| OBER manually audited primary cohort | 115 | 230 | External detector and policy evaluation |
| Setting | Qwen | Gemma |
|---|---|---|
| Routing features | 6,144 | 3,840 |
| Feature statistics | Per-feature GQA-train mean/std | Per-feature GQA-train mean/std |
| Loss and optimizer | BCE; AdamW | BCE; AdamW |
| Learning rate; batch size | ; 512 | ; 512 |
| Maximum epochs; patience | 30; 5 | 30; 5 |
| Epoch selection | Best validation ROC-AUC | Best validation ROC-AUC |
| Condition | Added instruction before the question |
|---|---|
| Standard | None |
| Generic | “Review the image carefully before answering.” |
| Target-aware | “Before answering, carefully verify whether the image actually contains any {target} . If it does not, explicitly state that the target is absent.” |
| Model | Cohort | ROC-AUC | Accuracy | Absence recall | Specificity | Trigger rate |
|---|---|---|---|---|---|---|
| Qwen | GQA held-out test | 0.9988 | 98.54 | 98.96 | 98.11 | 50.42 |
| Qwen | GQA audited review | 0.9995 | 98.75 | 99.17 | 98.33 | 50.42 |
| Qwen | OBER external | 0.8095 | 65.22 | 99.13 | 31.30 | 83.91 |
| Gemma | GQA held-out test | 0.9956 | 96.68 | 97.07 | 96.30 | 50.39 |
| Gemma | GQA audited review | 0.9998 | 99.17 | 100.00 | 98.33 | 50.83 |
| Gemma | OBER external | 0.7781 | 70.87 | 73.04 | 68.70 | 52.17 |
| Layer | ROC-AUC | Accuracy | Layer | ROC-AUC | Accuracy |
|---|---|---|---|---|---|
| 0 | 0.9948 | 97.22 | 24 | 0.9956 | 97.00 |
| 1 | 0.9980 | 98.24 | 25 | 0.9962 | 97.48 |
| 2 | 0.9986 | 98.54 | 26 | 0.9961 | 97.27 |
| 3 | 0.9987 | 98.69 | 27 | 0.9955 | 97.05 |
| 4 | 0.9987 | 98.61 | 28 | 0.9953 | 96.95 |
| 5 | 0.9987 | 98.61 | 29 | 0.9948 | 96.73 |
| Model | Threshold | Balanced acc. | Recall | Specificity | Trigger rate |
|---|---|---|---|---|---|
| Qwen | Frozen GQA | 65.22 | 99.13 | 31.30 | 83.91 |
| Qwen | OBER split-calibrated | 77.75 | 94.86 | 60.64 | 67.11 |
| Gemma | Frozen GQA | 70.87 | 73.04 | 68.70 | 52.17 |
| Gemma | OBER split-calibrated | 73.19 | 68.61 | 77.77 | 45.42 |
| Model | Dataset | Standard | Generic review | Target-aware review | Target-aware gain |
|---|---|---|---|---|---|
| Qwen | GQA | 66.75 | 67.67 | 89.00 | +22.25 |
| Qwen | OBER | 83.39 | 89.57 | 95.57 | +12.17 |
| Gemma | GQA | 80.00 | 80.75 | 93.42 | +13.42 |
| Gemma | OBER | 93.91 | 94.09 | 95.30 | +1.39 |
| Dataset | Prompt | Always acc. | Selective acc. | Review rate |
|---|---|---|---|---|
| GQA | Generic | 67.25% | 67.67% | 50.42% |
| GQA | Target-aware | 86.25% | 89.00% | 50.42% |
| OBER | Generic | 89.39% | 89.57% | 83.91% |
| OBER | Target-aware | 95.30% | 95.57% | 83.91% |
| Dataset | Target | Review prompt | Standard | Always | Selective |
|---|---|---|---|---|---|
| GQA | Absent | Generic | 37.17 | 39.00 | 39.00 |
| GQA | Present | Generic | 96.33 | 95.50 | 96.33 |
| GQA | Absent | Target-aware | 37.17 | 83.17 | 82.50 |
| GQA | Present | Target-aware | 96.33 | 89.33 | 95.50 |
| OBER | Absent | Generic | 66.78 | 79.83 | 79.83 |
| OBER | Present | Generic | 100.00 | 98.96 | 99.30 |
| Model | Dataset | Generic review | Target-aware review |
|---|---|---|---|
| Qwen | GQA | 0/10 (0.00%) | 5/10 (50.00%) |
| Qwen | OBER | 4/395 (1.01%) | 27/395 (6.84%) |
| Qwen | Combined | 4/405 (0.99%) | 32/405 (7.90%) |
| Gemma | GQA | 0/10 (0.00%) | 0/10 (0.00%) |
| Gemma | OBER | 10/178 (5.62%) | 15/178 (8.43%) |
| Gemma | Combined | 10/188 (5.32%) | 15/188 (7.98%) |
| Layer | ROC-AUC | Accuracy (%) | Layer | ROC-AUC | Accuracy (%) |
|---|---|---|---|---|---|
| 0 | 0.9006 | 83.19 | 24 | 0.9377 | 86.50 |
| 1 | 0.9360 | 87.50 | 25 | 0.9599 | 89.34 |
| 2 | 0.9836 | 94.33 | 26 | 0.9489 | 87.69 |
| 3 | 0.9454 | 89.08 | 27 | 0.9359 | 86.28 |
| 4 | 0.9618 | 90.74 | 28 | 0.9407 | 86.92 |
| 5 | 0.9559 | 89.58 | 29 | 0.9385 | 86.42 |