Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zero-shot methods reason over it as a single unit. We propose Atomic Visual Entailment (AVE), which decomposes the hypothesis into atomic facts, produces candidate predictions from both the full hypothesis and its facts using frozen vision-language models, and predicts the final label with a lightweight classifier trained only on how those candidates behave. We find that decomposition helps only when the hypothesis context is preserved: judging facts in isolation is worse than not decomposing at all. Full-hypothesis and atomic prediction make complementary errors, and learning which to trust recovers far more of that complementarity than majority voting, reaching 0.803 test accuracy on SNLI-VE without fine-tuning any vision-language model. AVE also localises the visual evidence behind its prediction without region-level supervision. These results suggest that learning which candidate prediction to trust can close much of the gap to fine-tuned systems, offering a practical alternative where fine-tuning a vision-language model directly would need more labelled data or compute than is available.
Figures & tables
Figure 1: AVE resolves prediction-pool disagreement that majority voting cannot. When candidates split evenly across labels, majority voting’s tiebreak selects a confidently wrong candidate, while learned selection identifies the correct one.
Figure 2: The AVE framework. Components marked with a snowflake are frozen; the learned selector is the only trained component. Given an image premise and a hypothesis, the Decomposer generates atomic facts, frozen VLMs produce candidate label–score pairs from the full hypothesis and from the atomic facts, the Selector predicts the final label and selects a representative candidate for grounding.
Atoms
Share
E
N
C
1
63.6
39.2
27.4
33.4
2
31.8
23.6
42.6
33.9
3
3.9
20.4
51.1
28.5
4+
0.7
25.0
49.2
25.8
Table 1: Atomic fact distribution on the test split. Share is the percentage of instances in each bucket; E, N and C are gold-label percentages within it.
Figure 3: (a) Instance-level disagreement with full-hypothesis prediction on test. Independent atomic prediction is excluded from the prediction pool. (b) Per-class recall on test.
Development
Test
Method
Acc.
F1
Acc.
F1
Best full-hypothesis
0.749
0.732
0.754
0.737
Best atomic
0.750
0.739
0.757
0.746
Best self-decomposed
0.762
0.757
0.765
0.760
AVE-MV intra (Qwen3)
0.752
0.744
0.756
0.748
AVE-MV intra (InternVL3)
0.761
0.751
0.768
0.757
Table 2: Best individual candidate per prediction strategy, majority voting, and learned selection. AVE-MV intra votes among the six candidates from one VLM; AVE-MV inter votes across all twelve candidates from both VLMs. Oracle reports the proportion of instances for which at least one of the K=12 candidates matches the gold label; the chance pool replaces those candidates with twelve independent random labels.
Figure 4: Grounding examples. For entailment, the phrase names visible evidence supporting the hypothesis; for contradiction, it names what the image shows instead. E and C denote entailment and contradiction.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Decomposer
Faith.
Compl.
Both
Meta-Llama-3.1-8B
95
88
88
Qwen3-8B
96
90
90
Qwen2.5-32B
98
96
96
Appendix
Table 5: Decomposition quality on 100 hypotheses sampled from the training split, annotated manually by an author. Faithfulness checks that no fact adds information absent from the hypothesis; completeness checks that all claims are covered. Counts out of 100.
Figure 5: Full-hypothesis accuracy on the development split for seven VLMs, using each model’s best prompt style and label derivation.
Feature variant
K
Features
Full-hypothesis
4
129
Simple atomic
4
129
Structured atomic
4
129
Full 12 methods
12
409
Appendix
Table 6: Feature variants compared during AVE-LS training. K is the number of candidates included.
HistGB
XGBoost
Feature variant
Acc.
F1
Acc.
F1
Full-hypothesis
0.790
0.787
0.788
0.785
Simple atomic
0.792
0.788
0.794
0.792
Structured atomic
0.793
0.790
0.793
0.791
Full 12 methods
0.798
0.796
0.801
0.798
Appendix
Table 7: Validation accuracy and macro-F1 for AVE-LS across classifier families and feature variants.
Parameter
Value
objective
multi:softprob
eval_metric
mlogloss
n_estimators
1,100
max_depth
3
learning_rate
0.025
subsample
0.88
Appendix
Table 8: XGBoost hyperparameters for AVE-LS, chosen by validation macro-F1.
Figure 6: Validation macro-F1 against selector training size for the four feature variants. Error bars show one standard deviation across three random training subsamples at each size.
Figure 7: Row-normalised confusion matrix for the best individual candidate on the test set.
Figure 8: Row-normalised confusion matrix for AVE-MV inter on the test set.
Figure 9: Row-normalised confusion matrix for AVE-LS on the test set.
Figure 10: Prediction accuracy by atom-count bucket for each method family on the test set.
Figure 11: Prompt used for hypothesis decomposition.
Vision-Language Models (VLMs) often achieve high performance on benchmarks while remaining "black boxes", yet they remain prone to hallucination or rely on superficial shortcuts. In this work, we propose a framework designed to enhance both performance and interpretability through De-compositional Evidence Grounding. Unlike monolithic inference approaches, our approach forces the model to decompose a global query into a sequence of atomic sub-questions, each requiring an explicit sub-answer and critically a localized evidence bounding box. By grounding intermediate logical steps (e.g. identifying a container, analyzing liquid properties, and assessing environmental context) in specific visual regions, we construct a structured reasoning path that mirrors human-like deduction. This allows the final answer to emerge as a logical consequence of verified visual facts rather than a statistical guess.
Eric Peh, Debaditya Roy, Basura Fernando
Institute of High-Performance Computing, Agency for Science, Technology and Research, Singapore · Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore · Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India +1
Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like "umbrella and no person" retrieve images containing both, even when concept detection is reliable. We trace this to an interface-level Bag-of-Concepts effect, where similarity scores approximate mean pooling of concept evidence regardless of operators. Although operator-dependent signals exist in text embeddings, they are too weak or misaligned to affect rankings. Fine-tuning does not reliably resolve this failure because the dominant bottleneck is how similarity aggregates evidence rather than what encoders represent. We propose factored inference, which separates evidence extraction from constraint execution, and introduce LCSE (Logic-Constrained Score Editing), a training-free method that executes constraints externally using concept scores from frozen encoders. We also introduce FACTOR-Bench, where LCSE achieves 85.5% accuracy versus 73.2% for the best fine-tuned baseline, 90.7% when applied to SigLIP 2, and improves NegBench COCO MCQ accuracy from 27.2% to 65.2% while preserving retrieval performance.
Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on localized patch-based embeddings that are insufficient to extract semantics in multi-step reasoning. We propose "Decompose, Look, and Reason" (DLR), a reinforced latent reasoning framework that dynamically decomposes queries into textual premises, extracts premise-conditioned continuous visual latents, and deduces answers through grounded rationales. We introduce a three-stage training pipeline and propose a novel Spherical Gaussian Latent Policy, to enable effective exploration in the latent space. Extensive experiments on vision-centric benchmarks show that DLR consistently outperforms strong baselines, including text-only, interleaved multimodal CoT, and latent reasoning methods, while providing superior stepwise interpretability.
Mengdan Zhu, Senhao Cheng, Liang Zhao
Emory University · University of Michigan, Ann Arbor